arXiv:2608.07923cs.CVcs.MM2026-08

无需训练,通过跨模态竞争消除误激活,提升音视频事件感知精度。

SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange

论文配图:SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange
图 1 · 摘自论文原文
  • 引入跨模态竞争机制,让标签共享证据并相互制约。
  • 在LLP数据集上,类型准确率提升7.45点,事件准确率提升5.04点。
  • 适用于多数据集迁移,配置固定不变,适合部署场景。

音视频事件感知(AVEP)需判断视频中事件何时发生、是否可听可见。现有无训练方法通过匹配冻结的音频与视觉特征与文本编码的事件名来查询新事件词汇。然而,相关标签共享证据,错误标签可能得分不低于正确标签,称为虚假共激活(FCA)。单一阈值无法同时剔除错误标签并保留所有正确标签。我们提出SCoPE,一种无训练框架,使所有查询标签在共享证据中竞争,且每模态引导另一模态的事件选择。推导出两标签情形下消除FCA的精确条件。在相同冻结的CLIP+CLAP主干下,相较于报告的AV²A结果,SCoPE在LLP上将Type@seg提升7.45点,Event@seg提升5.04点。同一固定配置可直接迁移至OV-AVEBench和VGGSound-AVEL100k,无需调整。

原文摘要 · Abstract (English)

Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training-free methods query new event vocabularies by matching frozen audio and visual features with text-encoded event names. However, related labels share evidence. An incorrect label can then score at least as high as a correct one. We call this a false co-activation (FCA). No scalar cutoff can reject the incorrect label while keeping every correct one. Class-specific thresholds may prevent that label from becoming a final prediction, but the FCA remains in the underlying score vector. We introduce SCoPE, a training-free framework in which all queried labels compete for shared evidence and each modality guides event selection in the other. We derive an exact condition for when this competition removes an FCA in a two-label fit. With identical frozen CLIP+CLAP backbones on LLP, SCoPE improves Type@seg by 7.45 points and Event@seg by 5.04 points compared with the reported AV$^2$A values. The same fixed configuration transfers unchanged to OV-AVEBench and VGGSound-AVEL100k.

音视频感知无训练跨模态竞争事件检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。