arXiv:2512.21094cs.CV2025-12中稿 · ICML被引 13

构建统一评估基准,全面测试文本生成音视频的对齐与真实感

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

  • 基于分类体系设计500个复杂提示,确保语义丰富性与物理合理性
  • 融合信号级指标与大模型主观评分,评估跨模态对齐与指令遵循能力
  • 揭示当前最强模型在音效真实性和细粒度同步上的显著不足

文本到音视频(T2AV)生成旨在从自然语言中合成时间连贯的视频与语义同步的音频,但其评估仍碎片化,常依赖单模态指标或范围有限的基准,无法捕捉跨模态对齐、指令遵循及复杂提示下的感知真实性。为此,我们提出T2AV-Compass,一个统一的T2AV系统综合评估基准,包含通过分类驱动流程构建的500个多样化且复杂的提示,以保证语义丰富性与物理合理性。此外,T2AV-Compass引入双层级评估框架,结合视频质量、音频质量与跨模态对齐的客观信号级指标,以及基于大语言模型作为评判者的主观评估协议,用于指令遵循与真实感判断。对11个代表性T2AV系统的广泛评估表明,即使最强模型也远未达到人类水平的真实感与跨模态一致性,在音效真实性、细粒度同步和指令遵循方面仍存在持续失败。这些结果揭示了未来模型的巨大改进空间,并凸显T2AV-Compass作为挑战性诊断测试平台的价值。

原文摘要 · Abstract (English)

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction following, and perceptual realism under complex prompts. To address this limitation, we present T2AV-Compass, a unified benchmark for comprehensive evaluation of T2AV systems, consisting of 500 diverse and complex prompts constructed via a taxonomy-driven pipeline to ensure semantic richness and physical plausibility. Besides, T2AV-Compass introduces a dual-level evaluation framework that integrates objective signal-level metrics for video quality, audio quality, and cross-modal alignment with a subjective MLLM-as-a-Judge protocol for instruction following and realism assessment. Extensive evaluation of 11 representative T2AVsystems reveals that even the strongest models fall substantially short of human-level realism and cross-modal consistency, with persistent failures in audio realism, fine-grained synchronization, instruction following, etc. These results indicate significant improvement room for future models and highlight the value of T2AV-Compass as a challenging and diagnostic testbed for advancing text-to-audio-video generation.

音视频生成多模态评估跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。