arXiv:2607.09091cs.CV2026-07中稿 · ECCV

用大模型把人类对音视频同步的主观判断转成客观评分,解决生成内容评估难题。

Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models

论文配图:Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models
图 1 · 摘自论文原文
  • 用人类对比排序生成失败样例构建数据集,再用Omni-LLM将相对评价转为连续绝对分值。
  • 提出新型优化方法,让评分能反映全局因果结构,与人类偏好高度一致。
  • 适合研究音视频生成、多模态评估或想摆脱人工标注依赖的开发者。

随着音视频生成模型发展为世界模拟器,跨模态同步成为评估生成内容中世界动态与因果一致性的重要指标。然而现有评估指标假设结构正确,仅将同步视为时间对齐,无法应对生成内容中的结构幻觉和非对称跨模态关系,当前亟需专家人工标注来评估同步性。这导致一个关键矛盾:人类评估依赖相对、有参考的比较,而自动化指标需无参考的绝对数值。本文通过将人类相对感知提炼为连续全局一致的度量,化解该矛盾。首先构建了基于成对人类标注的合成失败数据集SynthSync;其次利用配备连续潜在投影的Omni-LLM,将相对排名转化为连续绝对评分;最后提出实时值组相对策略优化($ R$-GRPO),通过列表式得分分布内化同步性的全局因果结构。实验表明,该度量在人类偏好对齐上达到最先进水平。我们以此建立标准化基准,推动音视频生成评估从低层次信号相关性跃升至视觉可解释的因果性层面。

原文摘要 · Abstract (English)

As audio-visual generative models evolve into world simulators, cross-modal synchronization stands as a critical proxy for assessing the consistency of world dynamics and causality in generated content. However, existing evaluation metrics presume structural correctness, reducing synchronization to mere temporal alignment. Consequently, they fail on generative outputs, especially when exhibiting structural hallucinations and asymmetric cross-modal relations, which currently \textbf{mandate expert human annotation to assess synchronization.} This dependency introduces a critical paradox: \emph{human evaluators rely on relative, reference-dependent comparisons, whereas automated metrics require reference-free, absolute scalars.} We resolve this paradox by proposing a framework that distills relative human perception into a continuous, globally consistent metric. First, we introduce SynthSync, a dataset of generative failures ranked via pairwise human annotations. Second, we adapt the Omni-LLM equipped with a continuous latent projection to translate relative human rankings into continuous absolute values. Third, we propose Real-Valued Group Relative Policy Optimization ($\mathbb{R}$-GRPO) to internalize the global causal structure of synchronization via listwise score distributions. Empirically, our metric achieves state-of-the-art human preference alignment. We leverage this estimator to establish a standardized benchmark, advancing AV-Gen assessment from low-level signal correlation to visually grounded causality.

音视频生成多模态评估大模型应用因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。