发现主流音频模型对相位信息不敏感,可能依赖频谱干扰而非真实相位编码。
Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models

- 用双耳掩蔽阈值差测试模型对微秒级相位结构的感知能力
- 仅专用空间音频模型达到接近理论基准的相位分辨能力
- 多数模型实则依赖宽带包络而非真实相位,易被语音信号干扰
近期空间自监督音频模型在定位任务上表现优异,引发对其是否编码微秒级双耳相位精细结构的疑问。本文提出基于双耳掩蔽水平差(BMLD)的心理声学评估基准,采用均衡抵消基线与GCC-PHAT正向对照,评估九个冻结音频模型,涵盖双耳自监督、单耳自监督及神经音频编码器。四个单耳负向对照模型均未产生有效BMLD,证实其双耳特异性。两个通用双耳自监督模型表现出极低相位敏感性,而专用空间双耳自监督模型达到接近解析基准的BMLD表现。渐进式物理消融实验表明,通用双耳模型依赖谱时干扰纹理,而非跨通道相位计算。语音中的高检测率反映其受宽带包络干扰,而非真实相位编码。
原文摘要 · Abstract (English)
Recent spatial self supervised audio models achieve high performance on localization tasks, raising questions about their encoding of microsecond interaural phase fine structures. We propose a psychoacoustic benchmark based on the binaural masking level difference to evaluate this. Using an equalization cancellation baseline and a GCC PHAT positive control we evaluate nine frozen audio models spanning binaural SSL, monaural SSL, and neural audio codecs. Four monaural negative controls yield zero BMLD confirming binaural specificity. Two general purpose binaural SSL models exhibit minimal phase sensitivity while dedicated binaural spatial SSL models achieve BMLD comparable to the analytical baseline. Progressive physical ablations show that general purpose binaural SSL models rely on spectro temporal interference textures rather than cross channel phase computation. High detection rates in speech reflect a confounding reliance on broadband envelopes rather than genuine phase encoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。