arXiv:2508.10360cs.SDeess.AS2025-08被引 2

为助听设备打造了公开可获取的声景识别数据集与轻量模型,实现低延迟实时识别。

A dataset and model for auditory scene recognition for hearing devices: AHEAD-DS and OpenYAMNet

  • 重构多个开源数据集形成AHEAD-DS,标签贴合助听场景需求
  • OpenYAMNet在14类声景上达86% mAP和93%准确率,适合边缘部署
  • 实测在旧手机上仅30ms/秒音频延迟,适合助听器等资源受限设备

声景识别对助听设备至关重要,但现有数据集普遍存在公开性差、标注不全或缺乏听觉相关标签的问题,限制了机器学习模型的系统性对比。为此,本文提出两方面解决方案:一是整合并优化多个开源数据集,构建专为助听设备设计的声景识别数据集AHEAD-DS;二是提出OpenYAMNet,一种面向边缘设备(如连接助听器的智能手机)的轻量级声音识别模型。AHEAD-DS提供标准化、公开可获取的数据集,标签与助听场景高度相关,便于模型比较。OpenYAMNet在AHEAD-DS测试集上对14类声景识别任务达到0.86的平均精度(mAP)和0.93的准确率。通过在安卓手机上部署,验证了其在边缘设备上的实时性能:即使在2018年款的谷歌Pixel 3上,模型加载延迟约50ms,每秒音频处理时间增加约30ms,满足低延迟要求。项目代码、数据与模型均已在GitHub公开。

原文摘要 · Abstract (English)

Scene recognition is important for hearing devices, however; this is challenging, in part because of the limitations of existing datasets. Datasets often lack public accessibility, completeness, or audiologically relevant labels, hindering systematic comparison of machine learning models. Deploying such models on resource-constrained edge devices presents another challenge.The proposed solution is two-fold, a repack and refinement of several open source datasets to create AHEAD-DS, a dataset designed for auditory scene recognition for hearing devices, and introduce OpenYAMNet, a sound recognition model. AHEAD-DS aims to provide a standardised, publicly available dataset with consistent labels relevant to hearing aids, facilitating model comparison. OpenYAMNet is designed for deployment on edge devices like smartphones connected to hearing devices, such as hearing aids and wireless earphones with hearing aid functionality, serving as a baseline model for sound-based scene recognition. OpenYAMNet achieved a mean average precision of 0.86 and accuracy of 0.93 on the testing set of AHEAD-DS across fourteen categories relevant to auditory scene recognition. Real-time sound-based scene recognition capabilities were demonstrated on edge devices by deploying OpenYAMNet to an Android smartphone. Even with a 2018 Google Pixel 3, a phone with modest specifications, the model processes audio with approximately 50ms of latency to load the model, and an approximate linear increase of 30ms per 1 second of audio. The project website with links to code, data, and models. https://github.com/Australian-Future-Hearing-Initiative

声景识别助听设备边缘计算轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。