综述音频视觉分割的进展与挑战,涵盖方法、数据集与未来方向。
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
- 系统梳理音视频融合的编码、对齐与解码架构设计。
- 对比多种训练范式在标准数据集上的性能表现差异。
- 适合关注多模态理解、弱监督学习的研究者参考。
音频-视觉分割(AVS)旨在通过结合视觉与音频模态,识别并分割视频中发声物体,是多模态感知的重要研究方向,支持细粒度的对象级理解。本文全面综述该领域,涵盖问题定义、基准数据集、评估指标及方法演进。分析了单模态与多模态编码架构、音视频融合策略、各类解码器设计,以及从完全监督到弱监督和无训练方法的主要训练范式。特别地,对主流方法在标准基准上的表现进行广泛比较,揭示不同架构选择、融合方式与训练范式对性能的影响。最后,指出当前挑战:时间建模有限、视觉模态偏差、复杂环境鲁棒性差、计算开销高,并提出未来方向:增强时序推理与多模态融合、利用基础模型提升泛化能力与少样本学习、通过自监督与弱监督减少标注依赖、引入高层推理以构建更智能的AVS系统。
原文摘要 · Abstract (English)
Audio-Visual Segmentation (AVS) aims to identify and segment sound-producing objects in videos by leveraging both visual and audio modalities. It has emerged as a significant research area in multimodal perception, enabling fine-grained object-level understanding. In this survey, we present a comprehensive overview of the AVS field, covering its problem formulation, benchmark datasets, evaluation metrics, and the progression of methodologies. We analyze a wide range of approaches, including architectures for unimodal and multimodal encoding, key strategies for audio-visual fusion, and various decoder designs. Furthermore, we examine major training paradigms, from fully supervised learning to weakly supervised and training-free methods. Notably, we provide an extensive comparison of AVS methods across standard benchmarks, highlighting the impact of different architectural choices, fusion strategies, and training paradigms on performance. Finally, we outline the current challenges, such as limited temporal modeling, modality bias toward vision, lack of robustness in complex environments, and high computational demands, and propose promising future directions, including improving temporal reasoning and multimodal fusion, leveraging foundation models for better generalization and few-shot learning, reducing reliance on labeled data through selfand weakly supervised learning, and incorporating higher-level reasoning for more intelligent AVS systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。