解决声音定位中的数据分布不均问题,提升真实场景下的定位精度。
Analytic Incremental Learning For Sound Source Localization With Imbalance Rectification
- 基于GCC-PHAT的增强方法缓解方向分布偏斜
- 新模型在SSLR上达89.0%准确率,误差5.3°
- 无需存储旧数据,适合持续学习场景
声音源定位(SSL)在受控环境下表现优异,但在实际应用中因双重不平衡挑战而性能下降:任务内不平衡源于长尾的方向到达(DoA)分布,任务间不平衡由跨任务偏差与重叠引起。这常导致灾难性遗忘,显著降低定位精度。为此,我们提出统一框架,包含两项关键创新:设计基于GCC-PHAT的数据增强(GDA)方法,利用峰值特征缓解任务内分布偏斜;提出解析式动态不平衡修正器(ADIR),结合任务自适应正则化,实现对任务间动态变化的解析更新。在SSLR基准上,该方法达到89.0%准确率、5.3°平均绝对误差和1.6的后向迁移分数,有效应对演化不平衡,且无需存储样本。
原文摘要 · Abstract (English)
Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance arising from long-tailed direction-of-arrival (DoA) distributions, and inter-task imbalance induced by cross-task skews and overlaps. These often lead to catastrophic forgetting, significantly degrading the localization accuracy. To mitigate these issues, we propose a unified framework with two key innovations. Specifically, we design a GCC-PHAT-based data augmentation (GDA) method that leverages peak characteristics to alleviate intra-task distribution skews. We also propose an Analytic dynamic imbalance rectifier (ADIR) with task-adaption regularization, which enables analytic updates that adapt to inter-task dynamics. On the SSLR benchmark, our proposal achieves state-of-the-art (SoTA) results of 89.0% accuracy, 5.3° mean absolute error, and 1.6 backward transfer, demonstrating robustness to evolving imbalances without exemplar storage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。