arXiv:2602.04853cs.CL2026-02ACL

用分解提示检测模型不确定,让大模型学会说‘我不知道’

Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"

  • 通过不同提示方式对比,发现模型意见不一致时更可能出错
  • 无需训练或检索,仅靠提示差异就能有效识别错误,提升准确率
  • 适合关注模型可靠性、防止幻觉的实用场景

大语言模型在闭卷问答中常无法识别自身知识边界,导致自信的幻觉。我们评估了三种等效提示策略:直接、辅助和增量式,在不同模型规模与多跳问答数据集上进行测试。结果表明,尽管前沿模型在分解提示下的准确率提升有限,但不同提示方式间的分歧仍高度预示潜在错误。由于事实知识稳定而幻觉具有随机性,跨提示一致性的信号能精准反映模型内部不确定性。我们据此提出一种无需训练、无需检索的拒答策略,实验显示该方法在多种设置下均优于标准不确定性基线,显著提升F1与AUROC指标。这证明分解提示可作为闭卷问答中模型可靠性的实用诊断工具。

原文摘要 · Abstract (English)

Large language models often struggle to recognize their knowledge limits in closed-book question answering, leading to confident hallucinations. While decomposed prompting is typically used to improve accuracy, we investigate its impact on reliability. We evaluate three task-equivalent prompting regimes: Direct, Assistive, and Incremental, across different model scales and multi-hop QA benchmarks. We find that although accuracy gains from decomposition diminish in frontier models, disagreements between prompting regimes remain highly indicative of potential errors. Because factual knowledge is typically stable while hallucinations are stochastic, cross-regime agreement provides a precise signal of internal uncertainty. We leverage this signal to implement a training-free abstention policy that requires no retrieval or fine-tuning. Our results show that disagreement-based abstention outperforms standard uncertainty baselines as an error detector, improving both F1 and AUROC across settings. This demonstrates that decomposition-based prompting can serve as a practical diagnostic probe for model reliability in closed-book QA.

模型可靠性提示工程拒绝回答幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。