解决音视频模态不同步问题,动态调整融合方式提升视频解析精度
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
- 提出动态交互机制,根据模态对齐程度自适应调整输入
- 在LLP数据集上显著优于现有方法,提升事件识别准确率
- 利用伪标签语义先验降低噪声干扰,适合多模态视频分析场景
音视频视频解析旨在通过弱标签对视频进行分类,同时识别事件是可见、可听或两者兼有的类型,并确定其时间边界。许多现有方法忽略了不同模态常存在非对齐现象,导致模态交互时引入额外噪声。本文提出一种学习非对齐知识的交互方法(LINK),通过动态调整各模态输入,均衡其在事件预测中的贡献。此外,利用伪标签的语义信息作为先验知识,以缓解其他模态带来的噪声。实验结果表明,该模型在LLP数据集上的表现优于现有方法。
原文摘要 · Abstract (English)
Audio-visual video parsing focuses on classifying videos through weak labels while identifying events as either visible, audible, or both, alongside their respective temporal boundaries. Many methods ignore that different modalities often lack alignment, thereby introducing extra noise during modal interaction. In this work, we introduce a Learning Interaction method for Non-aligned Knowledge (LINK), designed to equilibrate the contributions of distinct modalities by dynamically adjusting their input during event prediction. Additionally, we leverage the semantic information of pseudo-labels as a priori knowledge to mitigate noise from other modalities. Our experimental findings demonstrate that our model outperforms existing methods on the LLP dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。