用支持集适配CLAP模型,提升光纤声学识别在低样本场景下的性能。
CLAP-S: Support Set Based Adaptation for Downstream Fiber-optic Acoustic Recognition
- 基于支持集线性插值,融合微调知识与记忆检索实现跨域泛化
- 在实验室和真实场景数据集上均达到领先性能,尤其在低样本条件下
- 适合需要小样本适配的非传统声学传感器任务研究者参考
对比语言-音频预训练(CLAP)模型在多种声信号识别任务中表现卓越。光纤声学识别是其中关键下游任务,对环境感知具有重要意义。由于光纤传感器具有独特的频率响应和噪声特性,导致其部署环境属于领域特定、低样本且存在显著域偏移,因此将CLAP适配至该任务成为研究热点。为此,本文提出支持集适配方法CLAP-S,通过线性插值方式融合CLAP适配器与支持集,同时利用微调获得的隐式知识和记忆检索得到的显式知识,增强跨域泛化能力。实验表明,该方法在实验室采集的光纤版ESC-50数据集及真实世界光纤枪声-烟花数据集上均取得优异性能。本研究也为其他下游声学识别任务提供了宝贵启示。代码与数据集已公开于https://github.com/Jingchensun/clap-s。
原文摘要 · Abstract (English)
Contrastive Language-Audio Pretraining (CLAP) models have demonstrated unprecedented performance in various acoustic signal recognition tasks. Fiber-optic-based acoustic recognition is one of the most important downstream tasks and plays a significant role in environmental sensing. Adapting CLAP for fiber-optic acoustic recognition has become an active research area. As a non-conventional acoustic sensor, fiber-optic acoustic recognition presents a challenging, domain-specific, low-shot deployment environment with significant domain shifts due to unique frequency response and noise characteristics. To address these challenges, we propose a support-based adaptation method, CLAP-S, which linearly interpolates a CLAP Adapter with the Support Set, leveraging both implicit knowledge through fine-tuning and explicit knowledge retrieved from memory for cross-domain generalization. Experimental results show that our method delivers competitive performance on both laboratory-recorded fiber-optic ESC-50 datasets and a real-world fiber-optic gunshot-firework dataset. Our research also provides valuable insights for other downstream acoustic recognition tasks. The code and gunshot-firework dataset are available at https://github.com/Jingchensun/clap-s.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。