arXiv:2603.17025eess.AScs.AI2026-03中稿 · IEEE ICASSP 2026

用统一编码器让参考音与混响音对齐,提升声音定位精度。

Shared Representation Learning for Reference-Guided Targeted Sound Detection

论文配图:Shared Representation Learning for Reference-Guided Targeted Sound Detection
图 1 · 摘自论文原文
  • 共享表示空间统一处理参考音和混响音,减少结构复杂度
  • 在URBAN-SED数据集上达83.15%段级F1和95.17%准确率
  • 适合做声音分离与听觉注意力相关研究的读者

人类听众能通过选择性听觉注意从复杂声景中分离出目标声音,这启发了目标声音检测(TSD)的研究。该任务要求在提供目标声音参考音频的前提下,检测并定位混合音频中的目标声音。现有方法依赖为参考音频生成声学判别性条件嵌入向量,并与混合音频编码器配对,采用多任务学习联合优化。本文提出一种统一编码器架构,将参考音频与混合音频共同映射到共享表示空间,强化对齐效果的同时降低模型复杂度。遵循多任务训练范式,该方法显著优于先前方法,在URBAN-SED数据集上实现83.15%的段级F1分数和95.17%的整体准确率,建立新的基准。

原文摘要 · Abstract (English)

Human listeners exhibit the remarkable ability to segregate a desired sound from complex acoustic scenes through selective auditory attention, motivating the study of Targeted Sound Detection (TSD). The task requires detecting and localizing a target sound in a mixture when a reference audio of that sound is provided. Prior approaches, rely on generating a sound-discriminative conditional embedding vector for the reference and pairing it with a mixture encoder, jointly optimized with a multi-task learning approach. In this work, we propose a unified encoder architecture that processes both the reference and mixture audio within a shared representation space, promoting stronger alignment while reducing architectural complexity. This design choice not only simplifies the overall framework but also enhances generalization to unseen classes. Following the multi-task training paradigm, our method achieves substantial improvements over prior approaches, surpassing existing methods and establishing a new state-of-the-art benchmark for targeted sound detection, with a segment-level F1 score of 83.15% and an overall accuracy of 95.17% on the URBAN-SED dataset.

声音检测共享表示多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。