arXiv:2506.22920cs.AI2025-06ICML被引 6

通过自对弈游戏提升大模型推理时的理性能力

Improving Rationality in the Reasoning Process of Language Models through Self-playing Game

  • 设计自对弈博弈机制,让模型自我辩论与纠错
  • 在数学推理任务中,正确率提升显著,长链推理更稳定
  • 无需人工标注,适合需要自主反思的智能系统

大语言模型在数学、编程等任务中展现出较强的推理能力,但现有研究显示,即使最先进的模型也缺乏对其推理过程的真实理解。本文探索无监督条件下,通过自对弈机制提升模型推理的理性水平。我们设计了批判辨识游戏(Critic-Discernment Game, CDG),其中证明者首先给出问题解法,随后接受批评者的挑战——这些批评或旨在协助,或试图误导。证明者需在面对误导性评论时保持正确答案,同时根据建设性反馈修正错误。在涉及数学推理、逐步错误检测、自我修正及长链推理的任务上,实验表明CDG训练能显著增强已对齐大模型对其推理过程的理解能力。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated considerable reasoning abilities in various tasks such as mathematics and coding. However, recent studies indicate that even the best models lack true comprehension of their reasoning processes. In this paper, we explore how self-play can enhance the rationality of models in the reasoning process without supervision from humans or superior models. We design a Critic-Discernment Game(CDG) in which a prover first provides a solution to a given problem and is subsequently challenged by critiques of its solution. These critiques either aim to assist or mislead the prover. The objective of the prover is to maintain the correct answer when faced with misleading comments, while correcting errors in response to constructive feedback. Our experiments on tasks involving mathematical reasoning, stepwise error detection, self-correction, and long-chain reasoning demonstrate that CDG training can significantly improve the ability of well-aligned LLMs to comprehend their reasoning process.

推理增强自对弈模型反思

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。