首个专为音视频同步生成设计的综合性评测基准。
VABench: A Comprehensive Benchmark for Audio-Video Generation
- 构建三类任务与十五维评估体系,覆盖音画一致性与跨模态对齐。
- 涵盖七类内容场景,支持文本、图像到音视频的生成评测。
- 适合音视频生成研究者与模型开发者使用,推动领域标准建设。
近期视频生成技术进展显著,已能生成视觉吸引人且音画同步的视频。然而,现有视频生成评测基准在音视频协同生成方面评估不足,尤其针对音画同步输出的模型。为此,我们提出VABench,一个全面、多维度的评测框架,用于系统评估音视频同步生成能力。VABench包含三大任务类型:文本到音视频(T2AV)、图像到音视频(I2AV)和立体声音视频生成。其评估模块覆盖15个维度,包括文本-视频、文本-音频、视频-音频的配对相似性,音画同步性,口型-语音一致性,以及精心设计的音视频问答对等。评测覆盖七类内容:动物、人声、音乐、环境音、同步物理声、复杂场景和虚拟世界。通过系统分析与可视化结果,旨在建立音视频同步生成模型的新评估标准,推动该领域全面发展。
原文摘要 · Abstract (English)
Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks provide comprehensive metrics for visual quality, they lack convincing evaluations for audio-video generation, especially for models aiming to generate synchronized audio-video outputs. To address this gap, we introduce VABench, a comprehensive and multi-dimensional benchmark framework designed to systematically evaluate the capabilities of synchronous audio-video generation. VABench encompasses three primary task types: text-to-audio-video (T2AV), image-to-audio-video (I2AV), and stereo audio-video generation. It further establishes two major evaluation modules covering 15 dimensions. These dimensions specifically assess pairwise similarities (text-video, text-audio, video-audio), audio-video synchronization, lip-speech consistency, and carefully curated audio and video question-answering (QA) pairs, among others. Furthermore, VABench covers seven major content categories: animals, human sounds, music, environmental sounds, synchronous physical sounds, complex scenes, and virtual worlds. We provide a systematic analysis and visualization of the evaluation results, aiming to establish a new standard for assessing video generation models with synchronous audio capabilities and to promote the comprehensive advancement of the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。