arXiv:2511.13219cs.SDcs.AI2025-11被引 6

首个专为拟音音频生成设计的评测基准,解决现有数据与实际应用脱节问题。

FoleyBench: A Benchmark For Video-to-Audio Models

  • 构建包含5000个视频-音频-文本三元组的自动化数据集
  • 视频中声音与画面事件有因果关联,覆盖更广的拟音类别
  • 支持细粒度分析模型在音画对齐、时序同步等方面的表现

视频到音频生成(V2A)在影视后期、增强现实/虚拟现实和声音设计等领域日益重要,尤其适用于生成与屏幕动作同步的拟音效果。拟音要求音频在语义上与可见事件一致,并在时间上精确匹配其发生时机。然而,现有评估数据集与下游应用存在偏差:74%的过往数据集视频存在音画对应不佳的问题,且以语音和音乐为主,偏离拟音使用场景。为此,我们提出FoleyBench,首个专为拟音风格V2A设计的大规模评测基准。该数据集包含5000个(视频,真实音频,文本描述)三元组,每个片段均包含可见声源,且音频与画面事件具有因果关系。数据通过自动化可扩展流程从YouTube和Vimeo等真实网络视频中构建。相比以往数据集,FoleyBench在专为拟音设计的分类体系下,覆盖更丰富的声类。每段视频还标注了源复杂度、UCS/AudioSet类别和视频时长等元信息,支持对模型性能与失败模式的细粒度分析。我们对多个顶尖V2A模型进行了评测,涵盖音频质量、音画对齐、时序同步和音频-文本一致性。样本可访问:https://gclef-cmu.org/foleybench

原文摘要 · Abstract (English)

Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Foley requires generating audio that is both semantically aligned with visible events and temporally aligned with their timing. Yet, there is a mismatch between evaluation and downstream applications due to the absence of a benchmark tailored to Foley-style scenarios. We find that 74% of videos from past evaluation datasets have poor audio-visual correspondence. Moreover, they are dominated by speech and music, domains that lie outside the use case for Foley. To address this gap, we introduce FoleyBench, the first large-scale benchmark explicitly designed for Foley-style V2A evaluation. FoleyBench contains 5,000 (video, ground-truth audio, text caption) triplets, each featuring visible sound sources with audio causally tied to on-screen events. The dataset is built using an automated, scalable pipeline applied to in-the-wild internet videos from YouTube-based and Vimeo-based sources. Compared to past datasets, we show that videos from FoleyBench have stronger coverage of sound categories from a taxonomy specifically designed for Foley sound. Each clip is further labeled with metadata capturing source complexity, UCS/AudioSet category, and video length, enabling fine-grained analysis of model performance and failure modes. We benchmark several state-of-the-art V2A models, evaluating them on audio quality, audio-video alignment, temporal synchronization, and audio-text consistency. Samples are available at: https://gclef-cmu.org/foleybench

音视频生成拟音评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。