arXiv:2504.13055cs.CV2025-04NeurIPS被引 97

通过混合清晰与轻微失真图像,提升视觉推理模型的泛化能力。

NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation

  • 用干净和轻微失真的图像混合增强训练轨迹,引入感知多样性。
  • 在5个跨域评测中达到开源RL微调模型最佳表现,提升推理鲁棒性。
  • 无需额外训练成本,适配不同模型规模与数据量,易部署。

近期强化学习进展提升了视觉语言模型(VLMs)的推理能力,但如何增强策略探索以更好利用测试时计算资源仍待深入研究。同时,VLMs在不完美视觉感知下的表现持续受限,影响后续推理。本文提出NoisyRollout,一种简单有效的数据增强方法,通过混合来自清晰图像与适度失真图像的训练轨迹,注入感知多样性,促进更优策略探索,提升推理鲁棒性。采用噪声退火调度,训练初期强噪声促进探索,后期逐步减弱以保障稳定性。该方法无需额外训练开销,也不需修改强化学习目标。在2个不同训练数据集上的大量实验表明,NoisyRollout在5个跨域推理与感知基准上优于现有开源RL微调模型。此外,其在7B与32B模型、1K至6K数据规模及高斯噪声、旋转等不同图像增强类型下均表现稳定,验证了方法的通用性与可扩展性。

原文摘要 · Abstract (English)

Recent advances in reinforcement learning (RL) have strengthened the reasoning capabilities of vision-language models (VLMs). However, enhancing policy exploration to better scale test-time compute remains largely underexplored. In addition, VLMs continue to struggle with imperfect visual perception, which in turn affects the subsequent reasoning process. We introduce NoisyRollout, a simple yet effective data augmentation method that addresses these issues by mixing training trajectories from both clean and moderately distorted images. This approach injects perceptual diversity, encouraging better policy exploration and leading to more robust reasoning. A noise annealing schedule gradually reduces distortion strength, aiding exploration early in training while ensuring later stability. Crucially, our method is easy-to-adopt--requiring no additional training cost and no modifications to the RL objective. Extensive experiments on 2 distinct training datasets demonstrate that NoisyRollout achieves state-of-the-art performance among open-source RL-tuned models across 5 out-of-domain reasoning and perception benchmarks. Furthermore, we validate the effectiveness of NoisyRollout across model sizes (7B and 32B), data scales (from 1K to 6K) and image augmentation types (Gaussion noise and rotation), highlighting its generalizability and scalability.

视觉推理强化学习数据增强模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。