arXiv:2601.13752cs.AIcs.CL2026-01ACL被引 1

无需标注推理过程,用信念工程让大模型更高效更可信

Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering

  • 通过探测模型内部信念,用合成数据微调引导其行为
  • 在效率和忠实性任务上超越监督基线,训练成本更低
  • 适合希望低成本优化大模型推理能力的研究者

大型推理模型在复杂问题求解中表现卓越,但常存在计算冗余或推理不忠实的问题。现有行为调控方法多依赖强化学习或基于标准推理轨迹的微调,成本高且难扩展。本文发现大型推理模型具有内嵌的推理信念,可通过简单的逻辑探针捕捉。基于此,提出推理信念工程(RELIEF),通过将模型自我认知对齐至目标信念蓝图来塑造其行为。关键在于完全无需推理轨迹监督,仅需在合成的自省式问答对上微调,即可内化期望特质。大量实验表明,RELIEF在效率与忠实性任务上达到或超过监督基线,且训练成本更低。进一步分析验证了信念迁移可有效改变实际行为。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) have achieved remarkable success in complex problem-solving, yet they often suffer from computational redundancy or reasoning unfaithfulness. Current methods for shaping LRM behavior typically rely on reinforcement learning or fine-tuning with gold-standard reasoning traces, a paradigm that is both computationally expensive and difficult to scale. In this paper, we reveal that LRMs possess latent \textit{reasoning beliefs} that internally track their own reasoning traits, which can be captured through simple logit probing. Building upon this insight, we propose Reasoning Belief Engineering (RELIEF), a simple yet effective framework that shapes LRM behavior by aligning the model's self-concept with a target belief blueprint. Crucially, RELIEF completely bypasses the need for reasoning-trace supervision. It internalizes desired traits by fine-tuning on synthesized, self-reflective question-answering pairs that affirm the target belief. Extensive experiments on efficiency and faithfulness tasks demonstrate that RELIEF matches or outperforms behavior-supervised and preference-based baselines while requiring lower training costs. Further analysis validates that shifting a model's reasoning belief effectively shapes its actual behavior.

推理模型信念工程无监督调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。