让模型只定位用户指定的声音,像人一样在嘈杂环境中听准目标。
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios

- 用提示词引导注意力,只聚焦用户指定的目标声源。
- 在真实录音上准确估计目标方向和数量,优于现有方法。
- 适合需要精准声源定位的智能音箱、助听器等场景。
人类能在复杂声学场景中选择性关注特定声音并判断其方向,而当前基于深度学习的系统仍难以实现这种选择性定位。尽管声源定位(SSL)已取得显著进展,但多数方法会同时定位所有活跃声源,缺乏选择性;而目标声源分离(TSE)虽能通过多模态提示提取目标,却通常丢失用于精确定位的多通道相位信息。为此,本文提出提示引导的选择性目标声源定位任务,并设计SelectTSL端到端架构,仅定位用户指定的目标。核心创新在于引入提示引导的选择性注意力模块(PGSA),生成与提示相关的嵌入,指导通道间相位差(IPD)增强器优化原始相位线索,结合目标幅度信息联合估计到达方向(DoA)与目标声源基数(即目标数量)。该耦合设计有效聚焦于用户指定目标的空间特征,并可处理目标数量随时间变化的情况。在合成数据与真实录音上的大量实验表明,所提方法持续优于基线模型,且在真实声学环境中表现出强泛化能力。
原文摘要 · Abstract (English)
Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-based systems. Sound source localization (SSL) has achieved remarkable success with deep learning, yet most methods localize all active sources without selectivity. Conversely, target sound extraction (TSE) extracts sources using multimodal prompts but typically fails to preserve the multichannel spatial information required for accurate localization. To bridge this gap, we formulate the task of prompt-guided selective target sound localization and propose SelectTSL, an end-to-end architecture that localizes only the user-specified target in multi-source acoustic scenes. Specifically, we design a target-aware selective localization strategy that employs a Prompt-Guided Selective Attention Module (PGSA) to generate prompt-informed embeddings. These embeddings guide an inter-channel phase difference (IPD) enhancer to refine raw phase cues, fusing with target magnitudes to jointly estimate direction of arrival (DoA) and target-source cardinality, i.e., the number of target sound sources. This coupled design effectively focuses on the user-specified target spatial cues for selective localization and also handles time-varying numbers of target sources. Extensive experiments on both synthetic data and real-world recordings demonstrate that our proposed method consistently outperforms other baselines and exhibits robust generalization to real acoustic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。