arXiv:2511.15137cs.LGcs.AI2025-11被引 4

让大模型一边解题一边自检,提升推理可靠性。

From Solving to Verifying: A Unified Objective for Robust Reasoning in LLMs

  • 用统一损失函数同时优化解题和自我验证
  • 自检能力显著提升,推理表现基本不变
  • 适合需要高可靠性的推理任务应用

大语言模型的推理能力虽通过强化学习得到显著提升,但仍难以持续验证自身的推理过程。本文提出GRPO-Verif算法,通过统一损失函数联合优化解题生成与自我验证,并引入可调超参数控制验证信号权重。实验表明,该方法有效增强了模型的自检能力,同时保持了相近的推理性能。

原文摘要 · Abstract (English)

The reasoning capabilities of large language models (LLMs) have been significantly improved through reinforcement learning (RL). Nevertheless, LLMs still struggle to consistently verify their own reasoning traces. This raises the research question of how to enhance the self-verification ability of LLMs and whether such an ability can further improve reasoning performance. In this work, we propose GRPO-Verif, an algorithm that jointly optimizes solution generation and self-verification within a unified loss function, with an adjustable hyperparameter controlling the weight of the verification signal. Experimental results demonstrate that our method enhances self-verification capability while maintaining comparable performance in reasoning.

大模型推理自验证强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。