arXiv:2508.04566cs.CVcs.AI2025-08AAAI被引 6

用跨模态显著锚点提升弱监督音视频事件定位精度

CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization

  • 通过音频视觉一致性判断,自动找关键时间戳作为锚点
  • 在UnAV-100和ActivityNet1.3上达到当前最佳性能
  • 适合做弱监督音视频时序定位的研究者参考

密集音视频事件定位(DAVEL)任务旨在定位未剪辑视频中同时发生在音频与视觉模态的事件。本文研究一种更具挑战性的弱监督设置(W-DAVEL),仅提供视频级事件标签,事件时间边界未知。我们提出利用跨模态显著锚点,即在弱监督下预测可靠且音频视觉语义高度一致的时间戳。具体地,设计了互事件一致评估模块,通过比较音频与视觉预测类别差异生成一致得分;再通过全局视频与局部时间窗识别机制,在一致得分基础上确定跨模态锚点特征;融合后的锚点特征输入锚点引导的时间传播模块,增强原始时序特征的语义编码,从而提升弱监督下的定位效果。我们在UnAV-100和ActivityNet1.3数据集上建立基准,实验表明本方法性能达到当前最优。

原文摘要 · Abstract (English)

The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and more challenging weakly-supervised setting (W-DAVEL task), where only video-level event labels are provided and the temporal boundaries of each event are unknown. We address W-DAVEL by exploiting \textit{cross-modal salient anchors}, which are defined as reliable timestamps that are well predicted under weak supervision and exhibit highly consistent event semantics across audio and visual modalities. Specifically, we propose a \textit{Mutual Event Agreement Evaluation} module, which generates an agreement score by measuring the discrepancy between the predicted audio and visual event classes. Then, the agreement score is utilized in a \textit{Cross-modal Salient Anchor Identification} module, which identifies the audio and visual anchor features through global-video and local temporal window identification mechanisms. The anchor features after multimodal integration are fed into an \textit{Anchor-based Temporal Propagation} module to enhance event semantic encoding in the original temporal audio and visual features, facilitating better temporal localization under weak supervision. We establish benchmarks for W-DAVEL on both the UnAV-100 and ActivityNet1.3 datasets. Extensive experiments demonstrate that our method achieves state-of-the-art performance.

音视频定位弱监督跨模态锚点传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。