提出细粒度音视频理解新任务,实现声音区域级精准定位与描述。
RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding
- 构建双数据集支持细粒度音源理解,含声源掩码与逐帧文本描述。
- 在f-Music和f-Lifescene数据集上,模型实现音源分割与描述的最优性能。
- 适合关注音视频精细感知、多模态交互的研究者与应用开发者。
音频-视觉学习(AVL)是多模态学习与具身智能的核心任务,在场景理解与交互中具有重要意义。以往研究多聚焦于粗粒度任务(如音视频对应、声源定位、事件定位)。为提供更具体的场景感知细节,本文首次定义细粒度音频-视觉学习任务:区域感知声源理解(RA-SSU),目标是实现区域感知、帧级别、高质量的声源理解。为此,我们创新构建两个对应数据集:细粒度音乐(f-Music)与细粒度生活场景(f-Lifescene),分别包含3,976个样本(22类音乐场景,涵盖复杂乐器混音)与6,156个样本(61类生活声源)。我们提出SSUFormer模型,采用多模态输入输出架构,设计掩码协同模块(MCM)与分层提示专家混合(MoHE)模块,分别提升分割精度与描述丰富性。大量实验验证了任务可行性、数据集可用性及模型优越性,在声源理解基准上达到当前最佳(SOTA)性能。
原文摘要 · Abstract (English)
Audio-Visual Learning (AVL) is one fundamental task of multi-modality learning and embodied intelligence, displaying the vital role in scene understanding and interaction. However, previous researchers mostly focus on exploring downstream tasks from a coarse-grained perspective (e.g., audio-visual correspondence, sound source localization, and audio-visual event localization). Considering providing more specific scene perception details, we newly define a fine-grained Audio-Visual Learning task, termed Region-Aware Sound Source Understanding (RA-SSU), which aims to achieve region-aware, frame-level, and high-quality sound source understanding. To support this goal, we innovatively construct two corresponding datasets, i.e. fine-grained Music (f-Music) and fine-grained Lifescene (f-Lifescene), each containing annotated sound source masks and frame-by-frame textual descriptions. The f-Music dataset includes 3,976 samples across 22 scene types related to specific application scenarios, focusing on music scenes with complex instrument mixing. The f-Lifescene dataset contains 6,156 samples across 61 types representing diverse sounding objects in life scenarios. Moreover, we propose SSUFormer, a Sound-Source Understanding TransFormer benchmark that facilitates both the sound source segmentation and sound region description with a multi-modal input and multi-modal output architecture. Specifically, we design two modules for this framework, Mask Collaboration Module (MCM) and Mixture of Hierarchical-prompted Experts (MoHE), to respectively enhance the accuracy and enrich the elaboration of the sound source description. Extensive experiments are conducted on our two datasets to verify the feasibility of the task, evaluate the availability of the datasets, and demonstrate the superiority of the SSUFormer, which achieves SOTA performance on the Sound Source Understanding benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。