首个评估音视频空间对齐的基准,解决生成视频中声音位置不准的问题。
SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
- 构建音视频空间对齐数据集,按声音是否在画面内标注
- 提出新度量标准,量化音画空间一致性水平
- 对比联合与分阶段生成模型,揭示当前技术差距
本文针对多模态生成模型难以实现高质量音视频空间对齐的问题,提出全新的空间对齐音视频生成(SAVG)任务基准。构建了一个基于声音事件是否在画面内的音视频对齐数据集,并提出一种新的空间对齐评估指标。利用该数据集和指标,对两类基线方法进行评测:一类是联合音视频生成模型,另一类是先生成视频再生成对应音频的两阶段方法。实验结果表明,现有方法在视频质量、音频质量以及音画空间对齐程度上均与真实数据存在显著差距。
原文摘要 · Abstract (English)
This work addresses the lack of multimodal generative models capable of producing high-quality videos with spatially aligned audio. While recent advancements in generative models have been successful in video generation, they often overlook the spatial alignment between audio and visuals, which is essential for immersive experiences. To tackle this problem, we establish a new research direction in benchmarking the Spatially Aligned Audio-Video Generation (SAVG) task. We introduce a spatially aligned audio-visual dataset, whose audio and video data are curated based on whether sound events are onscreen or not. We also propose a new alignment metric that aims to evaluate the spatial alignment between audio and video. Then, using the dataset and metric, we benchmark two types of baseline methods: one is based on a joint audio-video generation model, and the other is a two-stage method that combines a video generation model and a video-to-audio generation model. Our experimental results demonstrate that gaps exist between the baseline methods and the ground truth in terms of video and audio quality, as well as spatial alignment between the two modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。