arXiv:2604.08540cs.CVcs.AI2026-04被引 11

构建多粒度评估基准,精准衡量文本生成音视频的语义准确性。

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

论文配图:AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
图 1 · 摘自论文原文
  • 设计任务驱动的多粒度评估框架,融合专用模型与多模态大模型。
  • 发现当前模型在语义可靠性上存在明显短板,尤其语音连贯性差。
  • 适合研究音视频生成、评估方法或内容可控性的研究人员参考。

文本到音视频(T2AV)生成正迅速成为媒体创作的核心接口,但其评估仍呈碎片化状态。现有基准大多孤立评估音频或视频,或依赖粗粒度的嵌入相似性,难以捕捉真实提示所需的细粒度联合正确性。我们提出 AVGen-Bench,一个面向11个现实场景的高质量提示任务驱动基准。为支持全面评估,引入多粒度评估框架,结合轻量级专用模型与多模态大语言模型(MLLMs),实现从感知质量到细粒度语义可控性的多层次评估。评估揭示:尽管音视频美学表现强劲,但语义可靠性严重不足,普遍存在文本渲染失败、语音连贯性差、物理推理错误及普遍的音乐音高控制失效问题。代码与资源详见 http://aka.ms/avgenbench。

原文摘要 · Abstract (English)

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity, failing to capture the fine-grained joint correctness required by realistic prompts. We introduce AVGen-Bench, a task-driven benchmark for T2AV generation featuring high-quality prompts across 11 real-world categories. To support comprehensive assessment, we propose a multi-granular evaluation framework that combines lightweight specialist models with Multimodal Large Language Models (MLLMs), enabling evaluation from perceptual quality to fine-grained semantic controllability. Our evaluation reveals a pronounced gap between strong audio-visual aesthetics and weak semantic reliability, including persistent failures in text rendering, speech coherence, physical reasoning, and a universal breakdown in musical pitch control. Code and benchmark resources are available at http://aka.ms/avgenbench.

音视频生成评估基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。