arXiv:2502.06664cs.SDcs.AI2025-02中稿 · International Conf…被引 1

评测深度音频表征在可穿戴设备中的实用性,发现BEATs模型表现最优。

Evaluation of Deep Audio Representations for Hearables

  • 构建首个针对可穿戴设备的音频表征评测基准DEAR,含1158段30秒音轨。
  • 在8项任务中,BEATs模型显著优于其他通用音频模型,尤其擅长环境感知。
  • 适合研究音频理解、智能听觉设备或语音交互的开发者和研究人员。

有效控制可穿戴听觉设备需要理解用户周围的声学环境。在声音场景的计算分析中,基础模型已成为生成高性能、鲁棒且多用途音频表征的前沿技术。本文提出并发布深度音频表征评测(DEAR),这是首个用于评估基础模型在捕捉可穿戴设备所需关键声学特性方面的有效性数据集与基准。数据集包含1,158段30秒长的音频轨道,通过将专有独白与高质量日常声学场景录音进行空间混合生成。该基准涵盖八项任务,评估音频场景的通用上下文、语音源及技术声学属性。通过对四个通用音频表征模型的评估,我们发现BEATs模型显著优于其他模型。这一优势凸显了在多样化音频数据上训练的模型在广泛听觉任务中的适用性,包括为可穿戴设备提供环境编码能力。DEAR数据集及相关代码已公开于https://dear-dataset.github.io。

原文摘要 · Abstract (English)

Effectively steering hearable devices requires understanding the acoustic environment around the user. In the computational analysis of sound scenes, foundation models have emerged as the state of the art to produce high-performance, robust, multi-purpose audio representations. We introduce and release Deep Evaluation of Audio Representations (DEAR), the first dataset and benchmark to evaluate the efficacy of foundation models in capturing essential acoustic properties for hearables. The dataset includes 1,158 audio tracks, each 30 seconds long, created by spatially mixing proprietary monologues with commercial, high-quality recordings of everyday acoustic scenes. Our benchmark encompasses eight tasks that assess the general context, speech sources, and technical acoustic properties of the audio scenes. Through our evaluation of four general-purpose audio representation models, we demonstrate that the BEATs model significantly surpasses its counterparts. This superiority underscores the advantage of models trained on diverse audio collections, confirming their applicability to a wide array of auditory tasks, including encoding the environment properties necessary for hearable steering. The DEAR dataset and associated code are available at https://dear-dataset.github.io.

音频表征可穿戴设备BEATs声学评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。