用音频模型教视觉模型,低资源下实现更准的声源定位与识别。
HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
- 通过分层跨模态蒸馏,让音频模型指导视听模型学习。
- 在DCASE数据集上相对基线提升21%-38%,媲美大模型效果。
- 适合资源有限但需高精度声源定位的智能监控、机器人场景。
本文提出HDA-SELD框架,结合分层跨模态蒸馏(HCMD)与多层级数据增强,解决低资源音视频声事件定位与检测(SELD)问题。以音频单模态模型为教师,通过输出响应和中间特征向量将知识传递给视听学生模型。通过随机混合多层网络特征并设计适配SELD任务的损失函数来增强学习。在DCASE 2023与2024挑战赛数据集上的大量实验表明,该方法显著提升视听SELD性能,整体指标相对基线提升21%-38%。值得注意的是,所提HDA-SELD在两项挑战赛中表现优于或相当甚至超过训练于更大数据集的教师模型,超越现有先进方法。
原文摘要 · Abstract (English)
This work presents HDA-SELD, a unified framework that combines hierarchical cross-modal distillation (HCMD) and multi-level data augmentation to address low-resource audio-visual (AV) sound event localization and detection (SELD). An audio-only SELD model acts as the teacher, transferring knowledge to an AV student model through both output responses and intermediate feature representations. To enhance learning, data augmentation is applied by mixing features randomly selected from multiple network layers and associated loss functions tailored to the SELD task. Extensive experiments on the DCASE 2023 and 2024 Challenge SELD datasets show that the proposed method significantly improves AV SELD performance, yielding relative gains of 21%-38% in the overall metric over the baselines. Notably, our proposed HDA-SELD achieves results comparable to or better than teacher models trained on much larger datasets, surpassing state-of-the-art methods on both DCASE 2023 and 2024 Challenge SELD tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。