arXiv:2503.13693cs.CV2025-03CVPR被引 3

无需训练即可识别新事件,动态调整判断标准。

Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds

  • 不依赖训练,用分数级融合保留音视频深层互动
  • 动态调节事件阈值,适应视频中事件分布变化
  • 适合处理未见事件,对弱监督和零样本场景有效

音频-视觉事件感知任务旨在跨模态(音频与视觉)进行事件的时空定位与分类。现有方法受限于训练数据中的词汇表,难以泛化到未见过的事件类别。此外,该任务标注成本高,需人工在多模态及时间片段上密集标注,限制了方法的可扩展性。当前先进模型忽略了事件分布随时间的变化,削弱了对视频动态演化的适应能力。以往方法采用晚融合策略结合音视频信息,虽简单但导致多模态交互信息大量丢失。为此,我们提出模型无关的训练免费方法 $ ext{AV}^2 ext{A}$,采用分数级融合以保留更丰富的多模态交互。$ ext{AV}^2 ext{A}$ 还引入帧内标签偏移算法,利用前一帧输入与预测结果动态调整后续帧的事件分布。我们首次构建了无需训练、开放词汇的基线,在零样本和弱监督设置下均显著优于朴素基线。实验表明,$ ext{AV}^2 ext{A}$ 在多个主流方法上均取得显著性能提升。

原文摘要 · Abstract (English)

In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the vocabulary available in their training data. This limitation significantly impedes their capacity to generalize to novel, unseen event categories. Furthermore, the annotation process for this task is labor-intensive, requiring extensive manual labeling across modalities and temporal segments, limiting the scalability of current methods. Current state-of-the-art models ignore the shifts in event distributions over time, reducing their ability to adjust to changing video dynamics. Additionally, previous methods rely on late fusion to combine audio and visual information. While straightforward, this approach results in a significant loss of multimodal interactions. To address these challenges, we propose Audio-Visual Adaptive Video Analysis ($\text{AV}^2\text{A}$), a model-agnostic approach that requires no further training and integrates a score-level fusion technique to retain richer multimodal interactions. $\text{AV}^2\text{A}$ also includes a within-video label shift algorithm, leveraging input video data and predictions from prior frames to dynamically adjust event distributions for subsequent frames. Moreover, we present the first training-free, open-vocabulary baseline for audio-visual event perception, demonstrating that $\text{AV}^2\text{A}$ achieves substantial improvements over naive training-free baselines. We demonstrate the effectiveness of $\text{AV}^2\text{A}$ on both zero-shot and weakly-supervised state-of-the-art methods, achieving notable improvements in performance metrics over existing approaches.

音视频感知零样本动态阈值多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。