arXiv:2604.25235cs.LGcs.CL2026-04被引 5

VLM做评分不可靠,任务不同不确定性差异大。

VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation

  • 用置信区间法将VLM评分转为带可靠范围的预测
  • 美学类任务区间窄(约40%),图表推理类达70%
  • 高排名相关性但分数不可靠,适合评估者注意

视觉语言模型(VLM)被广泛用于自动评判多模态系统,但其评分缺乏可靠性指示。本文通过置信区间方法,仅利用评分对应词元的概率分布,无需重新训练即可将点分转化为校准后的预测区间。这是首个对3个VLM评委在14类视觉任务中应用置信区间的系统分析。结果表明,评估不确定性强烈依赖于任务类型:美学与自然图像任务的区间覆盖约40%得分范围,而图表与数学推理任务扩展至约70%,形成多模态评估的定量可靠性图谱。此外,我们发现一种标准指标未捕捉的失效模式——排序与评分解耦:评委在排名上相关性高,但区间过宽且信息量低,能正确排序响应却无法给出可信绝对分数。最后,区间宽度主要由任务难度和标注质量决定,同一评委在同一方法下,在高质量多标注数据集上区间宽度缩小4.5倍。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly used as automated judges for multimodal systems, yet their scores provide no indication of reliability. We study this problem through conformal prediction, a distribution-free framework that converts a judge's point score into a calibrated prediction interval using only score-token log-probabilities, with no retraining. We present the first systematic analysis of conformal prediction for VLM-as-a-Judge across 3 judges and 14 visual task categories. Our results show that evaluation uncertainty is strongly task-dependent: intervals cover ~40% of the score range for aesthetics and natural images but expand to ~70% for chart and mathematical reasoning, yielding a quantitative reliability map for multimodal evaluation. We further identify a failure mode not captured by standard evaluation metrics, ranking-scoring decoupling, where judges achieve high ranking correlation while producing wide, uninformative intervals, correctly ordering responses but failing to assign reliable absolute scores. Finally, we show that interval width is driven primarily by task difficulty and annotation quality, i.e., the same judge and method yield 4.5x narrower intervals on a clean, multi-annotator captioning benchmark. Code: https://github.com/divake/VLM-Judge-Uncertainty

多模态评估置信区间VLM可靠性评分不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。