arXiv:2602.01623cs.CV2026-02被引 3

用全能大模型评估图文音生成效果,既准又可解释。

Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?

  • 用全能大模型统一分析图文音三模态输出
  • 在语义对齐任务上表现接近人工评价
  • 提供可理解的错误原因分析,适合优化生成

当前Sora 2和Veo 3等先进文本到视频生成模型能直接从文本提示生成高保真、音画同步的视频,标志着多模态生成的新里程碑。然而,对这种三模态输出的评估仍是未解难题。人工评估可靠但成本高且难扩展,传统自动指标如FVD、CLAP、ViCLIP仅关注单一模态对,难以处理复杂提示,且缺乏可解释性。全能大语言模型(omni-LLMs)因其天然处理音频、视频与文本的能力,支持丰富推理,提供可解释的思维链反馈,成为有前景的替代方案。我们提出Omni-Judge,研究omni-LLMs能否作为文本条件音频视频生成的人类对齐评判者。在九项感知与对齐指标中,Omni-Judge达到与传统指标相当的相关性,在音频-文本对齐、视频-文本对齐及音视频-文本一致性等语义挑战任务上表现优异。其在高帧率感知指标(如视频质量、音视频同步)上表现较差,因时间分辨率有限。Omni-Judge可生成可解释的说明,揭示语义或物理不一致,支持基于反馈的优化。研究凸显了omni-LLMs作为多模态生成统一评估者的潜力与当前局限。

原文摘要 · Abstract (English)

State-of-the-art text-to-video generation models such as Sora 2 and Veo 3 can now produce high-fidelity videos with synchronized audio directly from a textual prompt, marking a new milestone in multi-modal generation. However, evaluating such tri-modal outputs remains an unsolved challenge. Human evaluation is reliable but costly and difficult to scale, while traditional automatic metrics, such as FVD, CLAP, and ViCLIP, focus on isolated modality pairs, struggle with complex prompts, and provide limited interpretability. Omni-modal large language models (omni-LLMs) present a promising alternative: they naturally process audio, video, and text, support rich reasoning, and offer interpretable chain-of-thought feedback. Driven by this, we introduce Omni-Judge, a study assessing whether omni-LLMs can serve as human-aligned judges for text-conditioned audio-video generation. Across nine perceptual and alignment metrics, Omni-Judge achieves correlation comparable to traditional metrics and excels on semantically demanding tasks such as audio-text alignment, video-text alignment, and audio-video-text coherence. It underperforms on high-FPS perceptual metrics, including video quality and audio-video synchronization, due to limited temporal resolution. Omni-Judge provides interpretable explanations that expose semantic or physical inconsistencies, enabling practical downstream uses such as feedback-based refinement. Our findings highlight both the potential and current limitations of omni-LLMs as unified evaluators for multi-modal generation.

多模态评估大模型评测音视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。