arXiv:2409.07224cs.SDeess.AS2024-09被引 3

新方法让声音定位模型不断学新声音,还保护隐私。

Analytic Class Incremental Learning for Sound Source Localization with Privacy Protection

  • 用闭式解析解增量更新深度学习模型,避免遗忘旧知识。
  • 在公开数据集上达到90.9%定位准确率,优于现有方法。
  • 无需存储历史数据,适合智能家居等隐私敏感场景。

声源定位(SSL)技术广泛应用于监控和机器人领域。传统基于信号处理(SP)的方法在特定信号与噪声假设下可提供解析解,而近年基于深度学习(DL)的方法显著超越了它们。然而,其成功依赖大量训练数据和高算力,且通常需要大规模标注的空间数据,在适应新声音类别时表现不佳。为缓解上述挑战,我们提出一种新型类增量学习(CIL)方法——SSL-CIL,通过闭式解析解增量更新基于DL的SSL模型,有效避免灾难性遗忘。特别地,由于学习过程不回访任何历史数据(无实例存储),可保障数据隐私,更适用于智能家庭场景。在公开的SSLR数据集上的实证结果表明,本方法性能优越,定位准确率达90.9%,超过其他竞争方法。

原文摘要 · Abstract (English)

Sound Source Localization (SSL) enabling technology for applications such as surveillance and robotics. While traditional Signal Processing (SP)-based SSL methods provide analytic solutions under specific signal and noise assumptions, recent Deep Learning (DL)-based methods have significantly outperformed them. However, their success depends on extensive training data and substantial computational resources. Moreover, they often rely on large-scale annotated spatial data and may struggle when adapting to evolving sound classes. To mitigate these challenges, we propose a novel Class Incremental Learning (CIL) approach, termed SSL-CIL, which avoids serious accuracy degradation due to catastrophic forgetting by incrementally updating the DL-based SSL model through a closed-form analytic solution. In particular, data privacy is ensured since the learning process does not revisit any historical data (exemplar-free), which is more suitable for smart home scenarios. Empirical results in the public SSLR dataset demonstrate the superior performance of our proposal, achieving a localization accuracy of 90.9%, surpassing other competitive methods.

声音定位增量学习隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。