让视觉查询专注听声定位,提升音视频实例分割精度
Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
- 用跨注意力机制生成以声音为中心的查询,使其专注意不同声源
- 在AVISeg上实现+1.64 mAP、+0.6 HOTA、+2.06 FSLA提升
- 适合做音视频多模态目标分割的研究者和工程师
音视频实例分割(AVIS)需在视频序列中准确定位并追踪发声物体。现有方法因视觉偏见存在两大问题:统一加法融合导致查询无法区分不同声源;仅视觉训练目标使查询收敛至任意显著物体。本文提出基于交叉注意力的声源中心查询生成方法,使每个查询可选择性关注特定声源,并将声源先验带入视觉解码。同时引入声源感知序数计数(SAOC)损失,通过具有单调一致性约束的序数回归显式监督发声物体数量,防止训练中仅依赖视觉的偏差。在AVISeg基准上的实验表明,该方法实现一致性能提升:mAP提高1.64,HOTA提升0.6,FSLA提升2.06,验证了查询专业化与显式计数监督对精准音视频实例分割的关键作用。
原文摘要 · Abstract (English)
Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform additive fusion prevents queries from specializing to different sound sources, while visual-only training objectives allow queries to converge to arbitrary salient objects. We propose Audio-Centric Query Generation using cross-attention, enabling each query to selectively attend to distinct sound sources and carry sound-specific priors into visual decoding. Additionally, we introduce Sound-Aware Ordinal Counting (SAOC) loss that explicitly supervises sounding object numbers through ordinal regression with monotonic consistency constraints, preventing visual-only convergence during training. Experiments on AVISeg benchmark demonstrate consistent improvements: +1.64 mAP, +0.6 HOTA, and +2.06 FSLA, validating that query specialization and explicit counting supervision are crucial for accurate audiovisual instance segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。