arXiv:2604.05445cs.CLcs.AI2026-04ACL

让视觉语言奖励模型可解释,动态选关键维度并加权评分

Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling

  • 通过视觉感知门控机制,动态选择相关评价维度
  • 在321k数据集上表现优于现有开源模型,显著减少幻觉
  • 适合需要可解释性对齐的视觉语言模型研究者

视觉语言奖励建模面临困境:生成式方法可解释但慢,判别式方法高效却如黑盒。为此,我们提出VL-MDR框架,将评估动态分解为细粒度、可解释的维度。不同于输出单一标量,VL-MDR利用视觉感知门控机制识别相关维度(如幻觉、推理),并针对每个输入自适应加权。为此,我们构建了包含321,000个视觉语言偏好对的数据集,涵盖21个细粒度维度。大量实验表明,VL-MDR在VL-RewardBench等基准上持续优于现有开源奖励模型。此外,基于VL-MDR构建的偏好对能有效支持DPO对齐,缓解视觉幻觉并提升可靠性,提供一种可扩展的视觉语言模型对齐方案。

原文摘要 · Abstract (English)

Vision-language reward modeling faces a dilemma: generative approaches are interpretable but slow, while discriminative ones are efficient but act as opaque "black boxes." To bridge this gap, we propose VL-MDR (Vision-Language Multi-Dimensional Reward), a framework that dynamically decomposes evaluation into granular, interpretable dimensions. Instead of outputting a monolithic scalar, VL-MDR employs a visual-aware gating mechanism to identify relevant dimensions and adaptively weight them (e.g., Hallucination, Reasoning) for each specific input. To support this, we curate a dataset of 321k vision-language preference pairs annotated across 21 fine-grained dimensions. Extensive experiments show that VL-MDR consistently outperforms existing open-source reward models on benchmarks like VL-RewardBench. Furthermore, we show that VL-MDR-constructed preference pairs effectively enable DPO alignment to mitigate visual hallucinations and improve reliability, providing a scalable solution for VLM alignment.

奖励建模可解释性视觉语言对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。