arXiv:2509.13215eess.AS2025-09中稿 · paper: Workshop on…被引 1

用重要性加权对抗训练,让合成数据模型更好跟踪移动声源。

Importance-Weighted Domain Adaptation for Sound Source Tracking

  • 用RNN最后隐状态统一变长音频特征,解决序列长度不一问题。
  • 在真实环境中定位准确率提升18.7%,显著改善跨域性能。
  • 适合做移动声源追踪的合成数据训练,尤其关注方向覆盖差异。

近年来,深度学习极大推动了声源定位(SSL)的发展。然而,训练此类模型需大量标注数据,而真实录音因声源移动难以标注。尽管使用模拟混响响应(RIR)和噪声生成的合成数据是可行替代方案,但合成数据训练的模型在真实环境中存在领域偏移问题。无监督领域自适应(UDA)可通过对齐合成与真实数据域来缓解此问题,无需真实数据标签。现有方法多聚焦静态声源定位,未考虑声源追踪(SST)带来的两个挑战:第一,输入序列长度可变导致特征维度不一致;第二,合成与真实数据的方向覆盖不匹配,由部分重叠或批大小限制引起,称为方向多样性失配。为此,我们提出一种专为SST设计的新颖无监督域自适应方法,包含两项核心机制:利用循环神经网络的最终隐藏状态作为固定维度特征表示,以处理变长序列;进一步采用重要性加权对抗训练,优先对齐与真实域相似的合成样本。实验表明,该方法成功将合成数据训练的模型适配至真实环境,显著提升声源追踪性能。

原文摘要 · Abstract (English)

In recent years, deep learning has significantly advanced sound source localization (SSL). However, training such models requires large labeled datasets, and real recordings are costly to annotate in particular if sources move. While synthetic data using simulated room impulse responses (RIRs) and noise offers a practical alternative, models trained on synthetic data suffer from domain shift in real environments. Unsupervised domain adaptation (UDA) can address this by aligning synthetic and real domains without relying on labels from the latter. The few existing UDA approaches however focus on static SSL and do not account for the problem of sound source tracking (SST), which presents two specific domain adaptation challenges. First, variable-length input sequences create mismatches in feature dimensionality across domains. Second, the angular coverages of the synthetic and the real data may not be well aligned either due to partial domain overlap or due to batch size constraints, which we refer to as directional diversity mismatch. To address these, we propose a novel UDA approach tailored for SST based on two key features. We employ the final hidden state of a recurrent neural network as a fixed-dimensional feature representation to handle variable-length sequences. Further, we use importance-weighted adversarial training to tackle directional diversity mismatch by prioritizing synthetic samples similar to the real domain. Experimental results demonstrate that our approach successfully adapts synthetic-trained models to real environments, improving SST performance.

声源追踪领域自适应合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。