arXiv:2606.05544cs.SDeess.AS2026-06中稿 · Interspeech 2026

提出评估音频模型空间感知能力的新基准,发现现有模型存在系统性偏差。

Probing Spatial Structure in Pretrained Audio Representations

  • 构建可控实验框架,测试模型对声源与房间空间特征的编码能力。
  • 声源位置比房间属性更易解码,且不同模型响应差异大。
  • 适合研究音频表征、空间听觉或模型可解释性的学者使用。

预训练的空间音频编码器在感知任务中日益广泛应用,但其空间编码能力仍不明确。本文提出空间音频表征学习(SARL)基准,一个用于评估预训练音频模型空间信息的受控框架。SARL 检测声源级因素(方位角、仰角、距离、类别)和房间级因素(混响时间 RT60、体积、形状)。跨多种编码器的实验揭示三个模式:输入配置和训练范式影响空间编码;声源因素始终比房间因素更易解码;在受控扰动下的敏感性分析显示,模型对声源与房间变化的响应呈异质性。这些结果揭示了当前预训练音频表示中的系统性偏差。SARL 已开源,支持空间音频表征的可复现评估。

原文摘要 · Abstract (English)

Pretrained spatial audio encoders are increasingly used as general-purpose representations for perceptual tasks, yet their spatial encoding capabilities remain poorly understood. We introduce the Spatial Audio Representation Learning (SARL) benchmark, a controlled framework for evaluating spatial information in pretrained audio models. SARL probes source-level factors (azimuth, elevation, distance, class) and room-level factors (RT60, volume, shape). Experiments across diverse encoders reveal three patterns: input configuration and training paradigm shape spatial encoding; source factors are consistently easier to decode than room factors; and sensitivity analysis under controlled perturbations shows heterogeneous responses to source and room variation. These results reveal systematic biases in current pretrained audio representations. SARL is released as an open-source benchmark for reproducible evaluation of spatial audio representations.

音频表征空间感知模型评估可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。