arXiv:2508.04848cs.AI2025-08被引 7

RL微调后大模型在真实非理想场景下推理能力显著下降

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

  • 基于强化学习微调模型,测试其在三种非理想场景下的表现
  • 三类非理想场景中模型性能普遍大幅下降,暴露推理缺陷
  • 呼吁重视真实环境评估,适合关注模型鲁棒性的研究者阅读

强化学习(RL)已成为提升大语言模型(LLM)推理能力的关键技术,政策梯度算法因高效性与有效性主导后训练阶段。然而,现有基准多在理想化条件下评估模型推理,忽视了真实世界中的非理想场景。本文识别出三种具有实际意义的非理想场景:摘要推理、细粒度噪声抑制和上下文过滤,并受脑科学启发,提出人类在输入不完美时仍能可靠推理的研究方向。我们正式定义并评估这些挑战性场景,使用代表性政策梯度算法对三款大语言模型及一款先进视觉-语言模型(LVLM)进行强化学习微调,随后在八个公开数据集上测试其表现。结果表明,尽管强化学习微调提升了理想条件下的推理能力,但在所有三类非理想场景中性能均显著下降,暴露出先进推理能力的关键局限。虽提出场景特定修复方法,但当前方法仍无法有效解决这些推理缺陷。本工作揭示大模型推理能力常被高估,强调在非理想场景下评估模型的重要性。代码与数据将公开于XXXX。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has become a key technique for enhancing the reasoning abilities of large language models (LLMs), with policy-gradient algorithms dominating the post-training stage because of their efficiency and effectiveness. However, most existing benchmarks evaluate large-language-model reasoning under idealized settings, overlooking performance in realistic, non-ideal scenarios. We identify three representative non-ideal scenarios with practical relevance: summary inference, fine-grained noise suppression, and contextual filtering. We introduce a new research direction guided by brain-science findings that human reasoning remains reliable under imperfect inputs. We formally define and evaluate these challenging scenarios. We fine-tune three LLMs and a state-of-the-art large vision-language model (LVLM) using RL with a representative policy-gradient algorithm and then test their performance on eight public datasets. Our results reveal that while RL fine-tuning improves baseline reasoning under idealized settings, performance declines significantly across all three non-ideal scenarios, exposing critical limitations in advanced reasoning capabilities. Although we propose a scenario-specific remediation method, our results suggest current methods leave these reasoning deficits largely unresolved. This work highlights that the reasoning abilities of large models are often overstated and underscores the importance of evaluating models under non-ideal scenarios. The code and data will be released at XXXX.

大模型推理强化学习非理想场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。