arXiv:2510.10541cs.LGcs.AI2025-10ACL被引 1

现有强化学习评测基准无法真实反映模型泛化能力,存在严重误导。

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

  • 引入诊断工具与最优性能差距指标,量化训练/测试集表现差异
  • 发现训练集表现几乎与测试集持平,表明基准无法区分进步
  • 提出三原则:难度足够、评估平衡、分布鲁棒,指导未来评测设计

当前用于大语言模型强化学习的评测基准不足以衡量实际进展。尽管近期报告了强化学习在这些基准上的成绩提升,我们发现,在基准训练集上训练的模型性能几乎与直接在测试集上训练相当,说明这些基准无法可靠区分进一步的改进。为此,我们提出一套诊断工具和最优性能差距(OPG)度量,量化在训练集与测试集上训练的性能差异。通过压力测试进一步分析发现,尽管在基准上得分很高,现有强化学习方法在面对分布偏移、难度变化和反事实场景时仍表现不佳,而这些弱点正是当前基准未能揭示的。因此我们得出结论:现有基准不足以评估泛化能力,并提出三个核心设计原则以构建更可靠的评测:足够的难度、均衡的评估以及分布鲁棒性。

原文摘要 · Abstract (English)

Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs).Despite recent benchmark gains reported for RL, we find that training on these benchmarks' training sets achieves nearly the same performance as training directly on the test sets, suggesting that the benchmarks cannot reliably separate further progress.To study this phenomenon, we introduce a diagnostic suite and the Oracle Performance Gap (OPG) metric that quantifies the performance difference between training on the train split versus the test split of a benchmark. We further analyze this phenomenon with stress tests and find that, despite strong benchmark scores, existing RL methods struggle to generalize across distribution shifts, varying levels of difficulty, and counterfactual scenarios: shortcomings that current benchmarks fail to reveal.We conclude that current benchmarks are insufficient for evaluating generalization and propose three core principles for designing more faithful benchmarks: sufficient difficulty, balanced evaluation, and distributional robustness.

强化学习评测基准泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。