arXiv:2506.13888cs.CLcs.CV2025-06被引 3

用视觉专家和迭代训练提升多模态模型的验证能力。

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

  • 引入视觉专家与思维链,迭代优化偏好数据
  • 在多个基准上显著降低幻觉率,提升多模态推理
  • 适合做多模态对齐与奖励建模的研究者参考

基于可验证奖励的强化微调(RFT)已推动大语言模型发展,但在视觉-语言(VL)模型中仍处于探索阶段。视觉-语言奖励模型(VL-RM)是实现对齐的关键,但其训练面临两大挑战:一是自举困境,高质量训练数据依赖已有强模型,导致自我强化偏差;二是模态偏差与负例放大,当模型错误生成视觉属性时,会生成误导性偏好数据,进一步恶化训练。为此,我们提出一种基于视觉专家、思维链(CoT)推理和基于边距的拒绝采样的迭代训练框架。该方法能精炼偏好数据,增强结构化批判能力,并持续改进推理过程。在多个VL-RM基准上的实验表明,该方法在幻觉检测和多模态推理方面均表现更优,推动了视觉-语言模型与强化学习的对齐进展。

原文摘要 · Abstract (English)

Reinforcement Fine-Tuning (RFT) with verifiable rewards has advanced large language models but remains underexplored for Vision-Language (VL) models. The Vision-Language Reward Model (VL-RM) is key to aligning VL models by providing structured feedback, yet training effective VL-RMs faces two major challenges. First, the bootstrapping dilemma arises as high-quality training data depends on already strong VL models, creating a cycle where self-generated supervision reinforces existing biases. Second, modality bias and negative example amplification occur when VL models hallucinate incorrect visual attributes, leading to flawed preference data that further misguides training. To address these issues, we propose an iterative training framework leveraging vision experts, Chain-of-Thought (CoT) rationales, and Margin-based Rejection Sampling. Our approach refines preference datasets, enhances structured critiques, and iteratively improves reasoning. Experiments across VL-RM benchmarks demonstrate superior performance in hallucination detection and multimodal reasoning, advancing VL model alignment with reinforcement learning.

多模态对齐奖励建模幻觉检测强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。