arXiv:2606.22918cs.CVcs.GT2026-06被引 1

为每种视觉语言模型定制评估标准,提升视频物理一致性判断准确率。

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

论文配图:Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation
图 1 · 摘自论文原文
  • 针对不同VLM设计专属评估维度,避免统一标准带来的偏差。
  • 在16个VLM上测试,平均准确率提升约32%。
  • 可发现模型个体盲点,适合模型选型与缺陷分析。

视频生成与世界模型的物理一致性评估日益依赖视觉语言模型(VLM)作为自动化评判者,提供奖励信号、排序决策和数据筛选依据。然而,不同VLM因训练数据与架构差异,在内部表征中对物理现象的理解各不相同。若采用单一全局评估标准,会强制所有VLM遵循相同的评价轴线,忽略其实际感知能力。本文提出JudgeFit,一种迭代优化流程,用于发现每个VLM的专属评估分类体系。初始分类通过提示目标VLM对少量视频中的物理错误进行描述并聚类构建;随后通过诊断步骤:将VLM的各维度评分与人类常识判断校准,识别出不可靠或冗余的维度,并用LLM进行修复,直至收敛。我们进一步将其转化为基准,应用于16个涵盖8个模型家族的VLM。结果表明,经优化的分类体系在未见视频上的表现均优于全局基准,平均相对提升约32%。除整体准确性外,各模型的个性化评估画像揭示了总体排名无法预见的盲区,不同模型家族的可靠性模式差异显著。

原文摘要 · Abstract (English)

Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that provide reward signals, ranking decisions, and data-filtering criteria. Yet VLMs differ substantially in training data and architecture, encoding physical phenomena through distinct internal representations. A single global evaluation schema therefore gives every VLM the same axes of competence, regardless of what each can actually perceive. We propose JudgeFit, an iterative refinement procedure that discovers a per-VLM evaluation taxonomy. An initial taxonomy is constructed by prompting the target VLM to enumerate physics errors on a small set of videos and clustering the resulting descriptions. The taxonomy is then refined through a diagnostic step: we calibrate the VLM's per-dimension scores to human physical-commonsense ratings, diagnose which dimensions it scores unreliably or redundantly, and prompt an LLM to repair them, iterating until convergence. We further instantiate this procedure as a benchmark and apply it to 16 VLMs spanning eight model families. The refined taxonomy outperforms the global-schema baseline on held-out videos for every VLM tested, with a mean relative improvement of approximately 32%. Beyond aggregate accuracy, the per-VLM profiles expose model-specific blind spots that overall rankings cannot anticipate, with reliability patterns differing markedly across model families.

VLM评估物理一致性模型分析评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。