用分层框架实现自然语言查询的音频分离,高效且语义精准。
Text-Queried Audio Source Separation via Hierarchical Modeling
- 分两阶段分离:先全局语义对齐,再局部结构保持分离
- 仅用少量标注数据即达顶尖性能,复杂场景语义一致性强
- 支持任意文本指令解析,灵活操控音频内容
通过自然语言查询提取任意音频事件,是目标音频分离的新兴范式。现有方法主要面临两大挑战:在无监督单阶段架构中难以联合建模声学-文本对齐与语义感知分离;依赖大规模精确标注数据以弥补跨模态学习效率低下。为此,我们提出分层分解框架HSM-TSS,将任务解耦为全局-局部语义引导特征分离与结构保持的声学重建。采用双阶段机制,在分别对齐文本查询的全局语义特征空间和保留时频结构的AudioMAE局部特征空间中进行分离。首先通过预训练的Q-Audio架构实现全局语义分离;随后基于预测的全局特征,在AudioMAE特征上执行第二阶段局部语义分离并完成声学重建。此外,设计指令处理流程,将任意文本查询解析为结构化操作(提取或移除)及音频描述,实现灵活声音操控。该方法在数据高效训练下达到当前最优分离性能,同时在复杂听觉场景中保持优异的语义一致性。
原文摘要 · Abstract (English)
Target audio source separation with natural language queries presents a promising paradigm for extracting arbitrary audio events through arbitrary text descriptions. Existing methods mainly face two challenges, the difficulty in jointly modeling acoustic-textual alignment and semantic-aware separation within a blindly-learned single-stage architecture, and the reliance on large-scale accurately-labeled training data to compensate for inefficient cross-modal learning and separation. To address these challenges, we propose a hierarchical decomposition framework, HSM-TSS, that decouples the task into global-local semantic-guided feature separation and structure-preserving acoustic reconstruction. Our approach introduces a dual-stage mechanism for semantic separation, operating on distinct global and local semantic feature spaces. We first perform global-semantic separation through a global semantic feature space aligned with text queries. A Q-Audio architecture is employed to align audio and text modalities, serving as pretrained global-semantic encoders. Conditioned on the predicted global feature, we then perform the second-stage local-semantic separation on AudioMAE features that preserve time-frequency structures, followed by acoustic reconstruction. We also propose an instruction processing pipeline to parse arbitrary text queries into structured operations, extraction or removal, coupled with audio descriptions, enabling flexible sound manipulation. Our method achieves state-of-the-art separation performance with data-efficient training while maintaining superior semantic consistency with queries in complex auditory scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。