打造挑战性评测基准,检验视觉语言生成模型的判断能力。
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
- 构建涵盖多类任务的高质数据集,专为探测模型弱点设计。
- 顶尖模型如GPT-4o准确率仅65.4%,开源模型难超随机猜测。
- 揭示感知缺陷是主因,训练模型自判可提效14.7%。
视觉语言生成奖励模型(VL-GenRMs)在对齐与评估多模态AI系统中起关键作用,但其自身评估仍缺乏深入探索。现有方法多依赖传统视觉语言任务中的AI标注偏好标签,易引入偏差且难以有效挑战先进模型。为此,我们提出VL-RewardBench,一个涵盖通用多模态查询、视觉幻觉检测与复杂推理任务的综合性基准。通过结合样本筛选与人工验证的AI辅助标注流程,我们构建了1,250个高质量样本,专门用于探测VL-GenRMs的局限性。对16个主流大视觉语言模型的全面评估表明,该基准具备强挑战性:即使GPT-4o也仅达65.4%准确率,而领先开源模型Qwen2-VL-72B仍难突破随机猜测水平。重要的是,VL-RewardBench表现与使用最佳N采样的MMMU-Pro准确率高度相关(皮尔逊相关系数r > 0.9)。分析揭示三大改进启示:(i) 模型主要失败于基础视觉感知而非推理;(ii) 推理时扩展收益随模型容量差异显著;(iii) 训练模型自主判断可显著提升性能(7B模型提升14.7%准确率)。我们认为,该基准与实验洞察将推动VL-GenRMs的发展。
原文摘要 · Abstract (English)
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench's effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson's r $>$ 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。