arXiv:2605.20183cs.CV2026-05被引 3

首个全面评估多镜头音视频生成的基准,解决评价标准单一难题。

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

论文配图:MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
图 1 · 摘自论文原文
  • 构建涵盖四维度的动态评估框架,支持最多15镜头复杂叙事
  • 自适应修正镜头分割,主观评分采用逐实例标准,提升判断可靠性
  • 与人工评判高度一致(斯皮尔曼相关系数91.5%),适合模型对比研究

视频生成正从单镜头合成转向满足现实需求的多镜头音视频(MSAV)叙事。然而,现有评估基准在范围和数据多样性上受限,依赖僵化流程,难以系统可靠地评测前沿MSAV模型。为此,我们提出MSAVBench,首个面向多镜头音视频生成的综合性基准与自适应混合评估框架。该基准覆盖视频、音频、镜头数、参考内容四个关键维度,包含多样任务设置、最高达15镜头的复杂场景及非现实挑战情境。评估框架通过自适应自我修正的镜头分割机制、逐实例主观评分标准以及基于工具的证据提取,显著提升评估鲁棒性。实验表明,该框架与人类评判高度一致(斯皮尔曼等级相关系数达91.5%)。对19个先进闭源与开源模型的系统评估显示,当前系统仍难以实现导演级控制与细粒度音画同步;而模块化或代理式生成流水线展现出缩小闭源与开源模型差距的潜力。数据与代码已公开于https://github.com/ali-vilab/MSAVBench。

原文摘要 · Abstract (English)

Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However, evaluating such frontier models remains a fundamental challenge. Existing benchmarks are limited in scope and data diversity, and rely on rigid evaluation pipelines, preventing systematic and reliable assessment of modern MSAV models. To bridge these gaps, we introduce MSAVBench, the first comprehensive benchmark and adaptive hybrid evaluation framework for multi-shot audio-video generation. Our benchmark spans four key dimensions, video, audio, shot, and reference, covering diverse task settings, varying shot counts of up to 15, and challenging non-realistic scenarios. Our evaluation framework improves robustness through an adaptive self-correction mechanism for shot segmentation, instance-wise rubrics for subjective metrics, and tool-grounded evidence extraction for complex judgments. Furthermore, MSAVBench achieves high alignment with human judgments, reaching a Spearman rank correlation of 91.5%. Our systematic evaluation of 19 state-of-the-art closed- and open-source models shows that current systems still struggle with director-level control and fine-grained audio-visual synchronization, while modular or agentic generation pipelines offer a promising path toward narrowing the gap between open- and closed-source models. The benchmark data and evaluation code are publicly available at https://github.com/ali-vilab/MSAVBench.

音视频生成多镜头评估基准自适应评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。