arXiv:2508.18760cs.AIcs.CL2025-08AAAI被引 3

让大模型学会在无法回答时主动放弃,提升可信AI决策能力。

Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models

  • 通过认知监控与推理时干预,识别模型内部判断与外部回应的不一致。
  • 在未充分条件的问题上,使模型拒答率提升至92.3%,同时保持原有推理准确率。
  • 适用于需要高可靠性、避免错误回答的医疗、金融等关键领域应用。

大型推理模型(LRMs)在复杂推理任务中表现出色,但面对缺乏充分条件的数学题等固有不可答问题时,仍持续错误作答。本文系统分析了模型在面对不可答问题时的响应行为,发现其具备识别问题缺陷的认知能力,却未能正确表现拒答行为,暴露出内在认知与外在输出之间的偏差。为此,提出一种轻量级两阶段方法,结合认知监控与推理时干预,有效提升拒答率。实验表明,该方法显著改善模型在不可答问题上的拒答表现,平均拒答率达92.3%,同时维持原有推理性能,为构建更可信的AI系统提供新路径。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) have shown remarkable progress on complex reasoning tasks. However, some questions posed to LRMs are inherently unanswerable, such as math problems lacking sufficient conditions. We find that LRMs continually fail to provide appropriate abstentions when confronted with these unanswerable questions. In this paper, we systematically analyze, investigate, and resolve this issue for trustworthy AI. We first conduct a detailed analysis of the distinct response behaviors of LRMs when facing unanswerable questions. Then, we show that LRMs possess sufficient cognitive capabilities to recognize the flaws in these questions. However, they fail to exhibit appropriate abstention behavior, revealing a misalignment between their internal cognition and external response. Finally, to resolve this issue, we propose a lightweight, two-stage method that combines cognitive monitoring with inference-time intervention. Experimental results demonstrate that our method significantly improves the abstention rate while maintaining the overall reasoning performance.

大模型可信推理拒答机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。