arXiv:2505.01747eess.AScs.SD2025-05被引 3

让模型在推理时利用设备信息,提升声景分类精度。

Low-Complexity Acoustic Scene Classification with Device Information in the DCASE 2025 Challenge

  • 引入设备信息辅助分类,支持设备定制化建模。
  • 基线模型准确率达50.72%,设备微调后提升至51.89%。
  • 适合关注真实设备部署与低复杂度模型的研究者。

本文介绍了 DCASE 2025 挑战赛中的低复杂度声景分类任务(含基线系统),延续了以往(2022–2024)对低复杂度模型、数据效率和设备不匹配的关注。本年度关键变化在于:推理阶段提供录音设备信息,使模型可基于设备特性构建设备特定分类器,更贴近实际部署场景。训练集采用与对应 2024 年挑战赛相同的 25% 子集,不限制外部数据使用,凸显迁移学习的重要性。基线模型为设备无关模型,准确率为 50.72%,通过设备特定微调提升至 51.89%。共有 12 支队伍提交 31 项成果,其中 11 支队伍优于基线。最佳方案在评测集上较基线提升超 8 个百分点。

原文摘要 · Abstract (English)

This paper presents the Low-Complexity Acoustic Scene Classification with Device Information Task of the DCASE 2025 Challenge, along with its baseline system. Continuing the focus on low-complexity models, data efficiency, and device mismatch from previous editions (2022-2024), this year's task introduces a key change: recording device information is now provided at inference time. This enables the development of device-specific models that leverage device characteristics-reflecting real-world deployment scenarios in which a model is designed with awareness of the underlying hardware. The training set matches the 25% subset used in the corresponding DCASE 2024 challenge, with no restrictions on external data use, highlighting transfer learning as a central topic. The baseline achieves 50.72% accuracy with a device-agnostic model, improving to 51.89% when incorporating device-specific fine-tuning. The task attracted 31 submissions from 12 teams, with 11 teams outperforming the baseline. The top-performing submission achieved an accuracy gain of more than 8 percentage points over the baseline on the evaluation set.

声景分类设备信息低复杂度DCASE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。