arXiv:2603.16140cs.LG2026-03被引 3

真实噪声数据严重损害强化学习性能,现有方法无法弥补。

Noisy Data is Destructive to Reinforcement Learning with Verifiable Rewards

  • 重新验证数据后发现,所谓高噪声数据实则混有清洁数据
  • 真实噪声使模型在数学推理上比干净数据差8-10%
  • 适用于对数据质量敏感的场景,如文本转SQL任务

基于可验证奖励的强化学习(RLVR)推动了大语言模型在多个领域的进步。近期研究声称改进的RLVR算法能有效应对错误标注,性能接近清洁数据训练。本文指出这些结论无效——所谓100%噪声数据实际混有清洁样本。通过严格重验证流程修正数据后,我们发现噪声对RLVR具有破坏性:现有算法改进无法缓解噪声影响,性能与基础GRPO相当;在数学推理基准上,仅含错误标注的模型表现比清洁数据训练的模型低8-10%。此外,在真实文本转SQL任务中,人类标注错误导致准确率下降5-12%。结果表明当前RLVR方法尚无法补偿数据质量问题,高质量数据仍至关重要。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has driven recent capability advances of large language models across various domains. Recent studies suggest that improved RLVR algorithms allow models to learn effectively from incorrect annotations, achieving performance comparable to learning from clean data. In this work, we show that these findings are invalid because the claimed 100% noisy training data is "contaminated" with clean data. After rectifying the dataset with a rigorous re-verification pipeline, we demonstrate that noise is destructive to RLVR. We show that existing RLVR algorithm improvements fail to mitigate the impact of noise, achieving similar performance to that of the basic GRPO. Furthermore, we find that the model trained on truly incorrect annotations performs 8-10% worse than the model trained on clean data across mathematical reasoning benchmarks. Finally, we show that these findings hold for real-world noise in Text2SQL tasks, where training on real-world, human annotation errors cause 5-12% lower accuracy than clean data. Our results show that current RLVR methods cannot yet compensate for poor data quality. High-quality data remains essential.

强化学习数据质量语言模型噪声鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。