arXiv:2601.00885cs.AI2026-01

让大模型自己质疑自己的推理,用反事实思考提升准确性和稳定性。

Counterfactual Self-Questioning for Stable Policy Optimization in Language Models

  • 模型自动生成反事实推理路径,挑战自身逻辑漏洞。
  • 在数学推理任务中准确率提升,小模型效果更显著。
  • 无需外部奖励模型,适合资源有限场景下的自我优化。

近期语言模型自改进研究显示,模型可通过反思、验证、辩论或自生成奖励来优化自身推理。然而,多数方法依赖外部批评者、学习到的奖励模型或集成采样,增加复杂性并引发训练不稳定。本文提出反事实自提问框架,仅用单个语言模型生成并评估自身推理的反事实批判。该方法先生成初始推理过程,提出针对潜在错误点的针对性问题,并生成替代推理轨迹以暴露错误假设或无效步骤。这些反事实轨迹提供结构化相对反馈,可直接用于策略优化,无需辅助模型。在多个数学推理基准上的实验表明,反事实自提问提升了准确率和训练稳定性,尤其对小模型效果显著,实现了仅靠内部生成监督的可扩展自我改进。

原文摘要 · Abstract (English)

Recent work on language model self-improvement shows that models can refine their own reasoning through reflection, verification, debate, or self-generated rewards. However, most existing approaches rely on external critics, learned reward models, or ensemble sampling, which increases complexity and training instability. We propose Counterfactual Self-Questioning, a framework in which a single language model generates and evaluates counterfactual critiques of its own reasoning. The method produces an initial reasoning trace, formulates targeted questions that challenge potential failure points, and generates alternative reasoning trajectories that expose incorrect assumptions or invalid steps. These counterfactual trajectories provide structured relative feedback that can be directly used for policy optimization without auxiliary models. Experiments on multiple mathematical reasoning benchmarks show that counterfactual self-questioning improves accuracy and training stability, particularly for smaller models, enabling scalable self-improvement using internally generated supervision alone.

自提问推理优化稳定训练语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。