arXiv:2507.04667cs.CVcs.AI2025-07ICCV被引 4

提出视频级音画定位新基准与模型,精准捕捉声音随时间的动态变化。

What's Making That Sound Right Now? Video-centric Audio-Visual Localization

  • 构建四类真实场景的视频级音画定位数据集
  • 新模型TAVLO实现高精度时序对齐,优于传统方法
  • 适合关注视频理解、多模态对齐的研究者

音画定位(AVL)旨在识别视觉场景中的发声源。现有研究多聚焦图像级音画关联,忽略时间动态性,且假设声源始终可见、仅涉及单一物体。为此,我们提出AVATAR——一个包含高分辨率时序信息的视频级音画定位基准,涵盖四种不同场景:单声音、混合声音、多实体和离屏场景,支持更全面的模型评估。同时,我们提出TAVLO,一种新型视频级音画定位模型,显式融合时间信息。实验表明,传统方法因依赖全局音频特征和帧级映射,难以追踪时序变化;而TAVLO通过高分辨率时序建模,实现稳健精确的音画对齐。本工作实证了时间动态在音画定位中的关键作用,并确立了视频级音画定位的新标准。

原文摘要 · Abstract (English)

Audio-Visual Localization (AVL) aims to identify sound-emitting sources within a visual scene. However, existing studies focus on image-level audio-visual associations, failing to capture temporal dynamics. Moreover, they assume simplified scenarios where sound sources are always visible and involve only a single object. To address these limitations, we propose AVATAR, a video-centric AVL benchmark that incorporates high-resolution temporal information. AVATAR introduces four distinct scenarios -- Single-sound, Mixed-sound, Multi-entity, and Off-screen -- enabling a more comprehensive evaluation of AVL models. Additionally, we present TAVLO, a novel video-centric AVL model that explicitly integrates temporal information. Experimental results show that conventional methods struggle to track temporal variations due to their reliance on global audio features and frame-level mappings. In contrast, TAVLO achieves robust and precise audio-visual alignment by leveraging high-resolution temporal modeling. Our work empirically demonstrates the importance of temporal dynamics in AVL and establishes a new standard for video-centric audio-visual localization.

音画定位视频理解多模态对齐时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。