arXiv:2602.12506cs.LG2026-02被引 10

RL微调的视觉语言模型易受文本干扰,导致推理不可靠。

On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

  • 通过强化学习微调提升视觉推理准确率,但牺牲了推理过程的可靠性。
  • 简单文本扰动使模型性能下降超30%,且一致性分析更暴露问题。
  • 开源模型鲁棒性差,闭源模型表现更好,反映当前开源训练方法不足。

强化学习(RL)微调已成为提升大语言模型在推理密集型任务中表现的关键技术,并正被扩展至视觉-语言模型(VLMs)。尽管经过RL微调的VLMs在视觉推理基准上表现提升,但仍对弱视觉定位、幻觉和过度依赖文本线索敏感。我们发现,简单的可控文本扰动(如误导性标题或错误的思维链追踪)会导致性能和置信度显著下降,尤其在跨开源多模态推理模型中评估思维链一致性时更为明显。相比之下,闭源模型表现出相似的失败模式,但具备更强的鲁棒性和推理一致性,表明差距源于当前开源RL微调方法的缺陷,而非任务本身局限。进一步分析揭示了准确率与可信度之间的权衡:微调虽提高基准准确率,却可能削弱思维链的可靠性及对上下文变化的鲁棒性。对抗性增强虽可改善鲁棒性,但无法阻止可信度漂移;引入可信度感知奖励可恢复答案与推理的一致性,但若与增强结合,训练易陷入捷径策略,鲁棒性仍难以保证。这些发现凸显了仅以准确率为评估标准的局限性,呼吁建立兼顾正确性、鲁棒性和视觉推理可信度的联合训练与评估协议。

原文摘要 · Abstract (English)

Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision-language models (VLMs). While RL-tuned VLMs improve on visual reasoning benchmarks, they remain vulnerable to weak visual grounding, hallucinations, and over-reliance on textual cues. We show that simple, controlled textual perturbations, including misleading captions or incorrect chain-of-thought (CoT) traces, cause substantial drops in robustness and confidence, and that these effects are more pronounced when CoT consistency is taken into account across open-source multimodal reasoning models. In contrast, closed models exhibit similar failure modes but maintain markedly greater robustness and reasoning consistency, suggesting that the gap reflects a shortcoming in current open-source RL finetuning rather than an inherent limitation of the task. To better understand these vulnerabilities, we further analyze RL finetuning dynamics and uncover an accuracy-faithfulness trade-off: finetuning raises benchmark accuracy, but can simultaneously erode the reliability of the accompanying CoT and its robustness to contextual shifts. Although adversarial augmentation improves robustness, it does not by itself prevent faithfulness drift. Incorporating a faithfulness-aware reward can restore alignment between answers and reasoning, but when paired with augmentation, training risks collapsing onto shortcut strategies and robustness remains elusive. Together, these findings highlight the limitations of accuracy-only evaluations and motivate training and assessment protocols that jointly emphasize correctness, robustness, and the faithfulness of visually grounded reasoning.

视觉语言模型强化学习推理鲁棒性思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。