arXiv:2601.18335cs.SDcs.AI2026-01中稿 · ICASSP26

解决声音定位中的数据分布不均问题,提升真实场景下的定位精度。

Analytic Incremental Learning For Sound Source Localization With Imbalance Rectification

  • 基于GCC-PHAT的增强方法缓解方向分布偏斜
  • 新模型在SSLR上达89.0%准确率,误差5.3°
  • 无需存储旧数据,适合持续学习场景

声音源定位(SSL)在受控环境下表现优异,但在实际应用中因双重不平衡挑战而性能下降:任务内不平衡源于长尾的方向到达(DoA)分布,任务间不平衡由跨任务偏差与重叠引起。这常导致灾难性遗忘,显著降低定位精度。为此,我们提出统一框架,包含两项关键创新:设计基于GCC-PHAT的数据增强(GDA)方法,利用峰值特征缓解任务内分布偏斜;提出解析式动态不平衡修正器(ADIR),结合任务自适应正则化,实现对任务间动态变化的解析更新。在SSLR基准上,该方法达到89.0%准确率、5.3°平均绝对误差和1.6的后向迁移分数,有效应对演化不平衡,且无需存储样本。

原文摘要 · Abstract (English)

Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance arising from long-tailed direction-of-arrival (DoA) distributions, and inter-task imbalance induced by cross-task skews and overlaps. These often lead to catastrophic forgetting, significantly degrading the localization accuracy. To mitigate these issues, we propose a unified framework with two key innovations. Specifically, we design a GCC-PHAT-based data augmentation (GDA) method that leverages peak characteristics to alleviate intra-task distribution skews. We also propose an Analytic dynamic imbalance rectifier (ADIR) with task-adaption regularization, which enables analytic updates that adapt to inter-task dynamics. On the SSLR benchmark, our proposal achieves state-of-the-art (SoTA) results of 89.0% accuracy, 5.3° mean absolute error, and 1.6 backward transfer, demonstrating robustness to evolving imbalances without exemplar storage.

声音定位增量学习不平衡学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。