发现强化学习训练中幻觉反而是提升视觉推理的关键
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models

- 用诱导幻觉的干扰方法测试模型依赖真实视觉信息的程度
- 在仅靠幻觉的条件下,强化学习仍显著提升模型表现
- 适合关注多模态模型训练机制与可信度的研究者
近期强化学习(RL)在大模型推理中的成功,推动了其在多模态大语言模型(MLLMs)后训练中的应用,以增强视觉推理能力。尽管多项研究报告性能提升,但尚不清楚RL训练是否真正使模型从视觉信息中学习。本文提出「幻觉作为提示」框架,通过引入诱导幻觉、模态特定的噪声干扰,移除或替换正确答案所需的关键视觉信息,迫使模型依赖幻觉进行推理。在训练与评估中同时应用此类干扰,可诊断RL训练动态并揭示数据集内在特性。在多个多模态推理基准上的实验证明,模型幻觉在RL训练中的作用远超以往认知:在纯幻觉诱导设置下,强化学习仍能显著提升推理性能,某些情况下甚至优于标准训练。这一发现挑战了现有对MLLM推理训练的假设,并推动更注重模态感知的强化学习训练设计。
原文摘要 · Abstract (English)
The recent success of reinforcement learning (RL) in large reasoning models has inspired the growing adoption of RL for post-training Multimodal Large Language Models (MLLMs) to enhance their visual reasoning capabilities. Although many studies have reported improved performance, it remains unclear whether RL training truly enables models to learn from visual information. In this work, we propose the Hallucination-as-Cue Framework, an analytical framework designed to investigate the effects of RL-based post-training on multimodal reasoning models from the perspective of model hallucination. Specifically, we introduce hallucination-inductive, modality-specific corruptions that remove or replace essential information required to derive correct answers, thereby forcing the model to reason by hallucination. By applying these corruptions during both training and evaluation, our framework provides a unique perspective for diagnosing RL training dynamics and understanding the intrinsic properties of datasets. Through extensive experiments and analyses across multiple multimodal reasoning benchmarks, we reveal that the role of model hallucination for RL-training is more significant than previously recognized. For instance, we find that RL post-training under purely hallucination-inductive settings can still significantly improve models' reasoning performance, and in some cases even outperform standard training. These findings challenge prevailing assumptions about MLLM reasoning training and motivate the development of more modality-aware RL-based training designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。