首个细粒度评估视频生成音频的基准,揭示模型在不同音效类别表现差异。
VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories
- 构建四类音频的多任务评测框架,涵盖音效、音乐、语音与歌唱。
- 11个前沿模型测试显示,语音与歌唱生成效果显著落后于音效。
- 引入13项无参考指标并验证人评一致性,适合研究多模态音频生成者。
视频到音频(V2A)生成对沉浸式多媒体体验至关重要,但其评估仍不充分。现有基准常以统一协议评估多种音频类型,忽视了不同音频类别的细粒度需求。为此,我们提出VidAudio-Bench,一个包含四项核心特性的多任务评测基准:(1) 广覆盖:涵盖声音效果、音乐、语音和歌唱四类代表性音频,在V2A与视频-文本到音频(VT2A)设置下进行评估;(2) 大规模评测:包含1,634个视频-文本对,并评测11个最先进的生成模型;(3) 全面度量:引入13项任务特定的无参考指标,系统评估音频质量、视频-音频一致性及文本-音频一致性;(4) 人类对齐:通过主观研究验证所有指标,证明其与人类偏好高度一致。实验结果表明,当前V2A模型在语音与歌唱生成上表现远低于音效。VT2A结果进一步揭示指令遵循与视觉引导之间的根本矛盾:更强的视觉条件提升视频-音频对齐,却常牺牲目标音频类别的生成准确性。这些发现使VidAudio-Bench成为诊断V2A系统的全面且可扩展框架,并为多模态音频生成提供了新洞见。
原文摘要 · Abstract (English)
Video-to-Audio (V2A) generation is essential for immersive multimedia experiences, yet its evaluation remains underexplored. Existing benchmarks typically assess diverse audio types under a unified protocol, overlooking the fine-grained requirements of distinct audio categories. To address this gap, we propose VidAudio-Bench, a multi-task benchmark for V2A evaluation with four key features: (1) Broad Coverage: It encompasses four representative audio categories - sound effects, music, speech, and singing - under both V2A and Video-Text-to-Audio (VT2A) settings. (2) Extensive Evaluation: It comprises 1,634 video-text pairs and benchmarks 11 state-of-the-art generation models. (3) Comprehensive Metrics: It introduces 13 task-specific, reference-free metrics to systematically assess audio quality, video-audio consistency, and text-audio consistency. (4) Human Alignment: It validates all metrics through subjective studies, demonstrating strong consistency with human preferences. Experimental results reveal that current V2A models perform poorly in speech and singing compared to sound effects. Our VT2A results further highlight a fundamental tension between instruction following and visually grounded generation: stronger visual conditioning improves video-audio alignment, but often at the cost of generating the intended audio category. These findings establish VidAudio-Bench as a comprehensive and scalable framework for diagnosing V2A systems and provide new insights into multimodal audio generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。