arXiv:2602.03516cs.LGcs.AI2026-02被引 1

让大模型从高质量错误答案中学习,提升推理能力。

Not All Negative Samples Are Equal: LLMs Learn Better from Plausible Reasoning

  • 用反向强化学习生成结构合理但答案错误的假负样本
  • 在7个数学推理任务上平均比现有方法提升2.03%
  • 可直接用于模型微调,适合想提升推理能力的研究者

从负样本中学习对提升大语言模型(LLM)推理能力具有巨大潜力,但现有方法将所有错误回答视为同等信息量,忽视了样本质量的重要性。为此,我们提出可解释负样本(PNS),通过反向强化学习(RL)训练专用模型,生成格式规范、结构连贯但最终结果错误的高质量负样本。该方法采用组合奖励机制,综合格式合规性、准确率反转、奖励模型评估和思维链评价,使生成结果几乎与正确解答无法区分。我们进一步验证了PNS作为即插即用数据源在三种骨干模型上的偏好优化效果,在七个数学推理基准上表现优异,平均优于其他负样本合成方法2.03%。

原文摘要 · Abstract (English)

Learning from negative samples holds great promise for improving Large Language Model (LLM) reasoning capability, yet existing methods treat all incorrect responses as equally informative, overlooking the crucial role of sample quality. To address this, we propose Plausible Negative Samples (PNS), a method that synthesizes high-quality negative samples exhibiting expected format and structural coherence while ultimately yielding incorrect answers. PNS trains a dedicated model via reverse reinforcement learning (RL) guided by a composite reward combining format compliance, accuracy inversion, reward model assessment, and chain-of-thought evaluation, generating responses nearly indistinguishable from correct solutions. We further validate PNS as a plug-and-play data source for preference optimization across three backbone models on seven mathematical reasoning benchmarks. Results demonstrate that PNS consistently outperforms other negative sample synthesis methods, achieving an average improvement of 2.03% over RL-trained models.

大模型推理增强负样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。