arXiv:2602.03307cs.SD2026-02被引 3

GRAM通过多通道自编码器学习真实环境中的空间音频表示,提升降噪与定位性能。

GRAM: Spatial general-purpose audio representations for real-world environments

  • 采用多通道掩码自编码器,捕捉真实环境中的空间声学特征。
  • 在仿真与实录数据上均超越现有模型,仅用少量数据即达最优。
  • 适合需要声源定位与复杂环境适应的语音处理任务。

音频基础模型能学习通用音频表征以支持多种下游任务。尽管其在单通道、无混响的音频片段上表现优异,但在存在混响与噪声的真实声学环境中表现有限。此外,多数模型忽略真实环境的空间维度,无法支持声源定位任务。为此,我们提出GRAM:一种通用的真实世界音频模型,采用多通道掩码自编码器高效学习空间音频表示。我们在高保真自然空间声学环境模拟数据及真实场景录音上标准化评估GRAM及其他音频基础模型,并发布两个互补的基准任务集:NatHEAR与RealSELD。结果表明,GRAM在NatHEAR和干净单通道版本HEAR上优于所有当前最先进的自监督音频基础模型,且训练数据仅为后者的几分之一。在仿真环境中,GRAM实现顶尖定位性能,并在RealSELD中有效泛化至真实录音。综上,GRAM显著推进了真实环境下鲁棒的空间音频基础模型发展。

原文摘要 · Abstract (English)

Audio foundation models learn general-purpose audio representations that facilitate a wide range of downstream tasks. While the performance of these models has greatly increased for conventional single-channel, dry audio clips, their success in real-world acoustic environments with reverberation and noise is limited. Furthermore, most audio foundation models ignore the spatial dimension of real-world acoustic environments, ruling out tasks involving sound localization. To address these limitations, we propose GRAM: a general-purpose real-world audio model that employs a multi-channel masked autoencoder to efficiently learn spatial audio representations. We evaluated GRAM and other audio foundation models in a standardized manner on high-quality simulations of naturalistic, spatial acoustic environments as well as recordings of real-world environments and release these two complementary benchmark task suites: NatHEAR and RealSELD. Our results demonstrate that GRAM outperforms all state-of-the-art self-supervised audio foundation models on NatHEAR and the clean, single-channel version HEAR, while using only a fraction of the training data. GRAM also shows state-of-the-art localization performance in simulated environments and generalizes efficiently to real-world recordings in RealSELD. Taken together, GRAM presents a significant advance toward robust spatial audio foundation models for real-world environments.

空间音频自编码器语音定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。