arXiv:2607.00726cs.CVcs.SD2026-07中稿 · Interspeech 2026被引 1

首个分离音视频时序与语义评估的基准,解决传统方法耦合问题。

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

论文配图:AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
图 1 · 摘自论文原文
  • 分离时序与语义评估,独立测试音视频同步能力。
  • 覆盖10场景5任务,含3269段视频、38390个样本,数据经人工验证。
  • 适合音视频对齐、多模态模型评估的研究者使用。

音视频特征提取是多模态理解与生成任务的基础。然而,现有评估协议存在维度偏差,通常仅关注语义匹配或时序偏移检测,且数据构建方式耦合,难以独立评估时序与语义一致性。本文提出 AV-SyncBench,首个完全分离时序与语义评估的音视频同步基准。基于真实场景视频,涵盖语音、音乐与声音,覆盖10种场景和5项挑战任务。数据通过自动过滤与人工验证,确保画面声源真实。基准包含3269个视频和38390个样本,并评估了五种代表性模型在对齐与下游任务中的特征质量。代码与数据集已公开。

原文摘要 · Abstract (English)

Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models exhibit dimensional bias, typically focusing on either semantic matching or temporal offset detection. Moreover, their data construction remains coupled, preventing independent assessment of temporal and semantic consistency. We propose AV-SyncBench, the first benchmark to fully separate temporal and semantic evaluation for audio-visual synchronization. Built from in-the-wild videos, it spans Voice, Music, and Sound across 10 scenarios and 5 challenge tasks. Data are automatically filtered and manually verified to ensure on-screen sound sources. The benchmark contains 3,269 videos and 38,390 samples, and we evaluate five representative models to quantify feature quality for alignment and downstream tasks. The code and dataset are available at: https://fgt7t6g.github.io/AV-SyncBench.

音视频同步多模态评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。