Event date · · Chain-of-Self-Questioning

When Should LLMs Abstain? Chain-of-Self-Questioning for Selective Risk Control

FACT STATEMENT

A paper introduces Chain-of-Self-Questioning (CoSQ), a prompt-only framework that makes answer commitment conditional on an explicit assessment of the information required to answer a question. Three CoSQ variants were evaluated under seventeen conditions on the 817-item TruthfulQA multiple-choice validation set using eleven open-weight and hosted model families. In the final balanced-option protocol, Grounded-CoSQ at τ=0.90 reduces the mean unconditional wrong-commitment rate from 13.1% under chain-of-thought prompting to 8.9%, a 32.1% relative reduction, while increasing answered accuracy from 86.9% to 89.7% and answering 87.6% of questions. Both improvements hold for all eleven models and at every evaluated threshold. Critical-CoSQ and Adaptive-CoSQ provide neighboring operating points with 88.6% and 86.5% coverage, respectively, while remaining more reliable than the baseline. A secondary Natural Questions Short-Answer evaluation provides convergent open-form evidence.

What happened

Chain-of-Self-Questioning (CoSQ) is a prompt-only framework that conditions answer commitment on an explicit assessment of the information required to answer a question. Evaluated on TruthfulQA with eleven model families, Grounded-CoSQ at τ=0.90 reduces wrong commitments by 32.1% relative to chain-of-thought prompting while improving answered accuracy and maintaining high coverage. The approach enables tunable answer-or-abstain decisions without model fine-tuning.

Technical significance

CoSQ uses self-generated questions to assess whether the model has sufficient information before committing to an answer. The prompt-only design avoids architectural changes and can be applied across model families. The reported improvements are consistent across all eleven models and thresholds, suggesting robustness. The secondary Natural Questions evaluation indicates potential generalization beyond multiple-choice formats.

Industry impact

Selective risk control is becoming a practical requirement for deploying LLMs in high-stakes applications. Prompt-only abstention mechanisms like CoSQ offer a low-cost way to reduce unsupported answers without retraining or fine-tuning. The ability to tune coverage via threshold τ provides an operational dial for balancing answer rate and reliability.

Decision value

CoSQ can reduce the risk of incorrect answers in customer-facing and enterprise AI systems, potentially lowering liability and improving user trust. The prompt-only nature means deployment requires no additional training infrastructure, making it cost-effective for existing LLM-based products.

What to watch

Next signals include replication on open-ended generation tasks, integration with retrieval-augmented generation, and adoption in production systems where abstention is preferable to incorrect answers. Further research may explore combining CoSQ with confidence calibration or uncertainty estimation methods.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.