让声音场景模型能零样本适应新类别,无需重新训练。
Lightweight and Generalizable Acoustic Scene Representations via Contrastive Fine-Tuning and Distillation
- 用对比学习构建语义结构化的声景嵌入空间
- 在未见类别上实现少样本快速适应,闭集性能不降
- 适合边缘设备上需动态扩展类别的声音识别场景
边缘设备上的声音场景分类(ASC)模型通常受限于固定类别假设,难以适应现实应用中新增或细化的声音类别。本文提出 ContrastASC,通过构建保留场景间语义关系的嵌入空间,使模型能够无需重新训练即可适应未见类别。该方法结合预训练模型的监督对比微调与对比表示蒸馏,将结构化知识迁移到轻量级学生模型中。实验表明,ContrastASC 在保持强闭集性能的同时,显著提升了对未见类别的少样本适应能力。
原文摘要 · Abstract (English)
Acoustic scene classification (ASC) models on edge devices typically operate under fixed class assumptions, lacking the transferability needed for real-world applications that require adaptation to new or refined acoustic categories. We propose ContrastASC, which learns generalizable acoustic scene representations by structuring the embedding space to preserve semantic relationships between scenes, enabling adaptation to unseen categories without retraining. Our approach combines supervised contrastive fine-tuning of pre-trained models with contrastive representation distillation to transfer this structured knowledge to compact student models. Our evaluation shows that ContrastASC demonstrates improved few-shot adaptation to unseen categories while maintaining strong closed-set performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。