arXiv:2506.00934cs.SDcs.AI2025-06被引 3

GRAM通过多通道掩码自编码器学习真实环境下的空间音频表示,提升音源定位与泛化能力。

GRAM: Spatial general-purpose audio representation models for real-world applications

  • 采用多通道掩码自编码器,捕捉真实场景中的空间音频特征。
  • 在模拟与真实环境中均实现领先音源定位性能,且训练数据量更少。
  • 适合需要空间感知的语音识别、智能设备和声学监控等实际应用。

音频基础模型可学习通用音频表征,广泛支持下游任务。尽管其在单通道干音频上表现优异,但在存在混响和噪声的真实声学环境中性能受限,且多数模型忽略真实环境的空间维度,难以支持音源定位任务。为此,我们提出GRAM:一种通用型真实世界音频模型,采用多通道掩码自编码器高效学习空间音频表征。我们在高质量自然空间声学环境仿真及真实场景录音上,以标准化方式评估GRAM与其他音频基础模型,并发布两个互补的基准测试套件:NatHEAR与RealSELD。结果表明,GRAM在NatHEAR和干净单通道版本HEAR上优于所有现有自监督音频基础模型,且仅使用极少训练数据;在模拟环境中实现最先进的定位性能,并高效泛化至真实录音的RealSELD。总体而言,GRAM显著推进了面向真实环境的鲁棒空间音频基础模型的发展。

原文摘要 · Abstract (English)

Audio foundation models learn general-purpose audio representations that facilitate a wide range of downstream tasks. While the performance of these models has greatly increased for conventional single-channel, dry audio clips, their success in real-world acoustic environments with reverberation and noise is limited. Furthermore, most audio foundation models ignore the spatial dimension of real-world acoustic environments, ruling out tasks involving sound localization. To address these limitations, we propose GRAM: a general-purpose real-world audio model that employs a multi-channel masked autoencoder to efficiently learn spatial audio representations. We evaluated GRAM and other audio foundation models in a standardized manner on high-quality simulations of naturalistic, spatial acoustic environments as well as recordings of real-world environments and release these two complementary benchmark task suites: NatHEAR and RealSELD. Our results demonstrate that GRAM outperforms all state-of-the-art self-supervised audio foundation models on NatHEAR and the clean, single-channel version HEAR, while using only a fraction of the training data. GRAM also shows state-of-the-art localization performance in simulated environments and generalizes efficiently to real-world recordings in RealSELD. Taken together, GRAM presents a significant advance toward robust spatial audio foundation models for real-world environments.

空间音频自监督学习音源定位音频模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。