arXiv:2607.08800cs.SD2026-07

通过注入随机噪声,让普通语音大模型实现零样本立体声定位。

Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering

论文配图:Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering
图 1 · 摘自论文原文
  • 用双分支编码器分别处理左右声道,不依赖专用空间模块。
  • 在0.5音量偏移下仍达97.2%定位准确率,支持未见过的空间配置。
  • 无需微调即可泛化到新场景,适合想提升语音模型空间感知的开发者。

多模态大语言模型具备出色的语义音频理解能力,但因依赖单通道音频表示,仍缺乏空间感知能力。现有空间音频方法多依赖复杂的房间模拟和定制训练的几何感知立体编码器,限制了其可及性与泛化性。本文提出Dual-BEATs架构,将左右声道分别输入两个相同的语义编码器,替代专用空间模块。为突破内部归一化层抹除声道间差异的瓶颈,我们在编码前注入静态、不相关的抖动噪声,建立宏观方差基线,使空间信息‘绕过’归一化层。在三元方向分类任务(左、中、右)上,经抖动处理的模型达到高达97.2%的定位准确率,即使在细微0.5倍偏移幅度下依然有效,并展现出对全新空间配置的强零样本泛化能力。结果表明,在恰当的声学正则化下,标准多模态模型天然具备通用立体声理解能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (LLMs) have remarkable semantic audio understanding, yet they remain "spatially agnostic" due to their reliance on mono-channel audio representations. Currently, spatial audio perception methods mainly focus on complex room simulations and custom-trained, geometry-aware stereo encoders, which limits their accessibility and generalizability. In this paper, we introduce the Dual-BEATs architecture, in which the left and right audio channels are routed independently through two identical semantic encoders as an alternative to specialized spatial modules. To circumvent the architectural bottleneck where internal normalization otherwise erases the inter-channel variance of stereo audio, we inject a static, uncorrelated dithering noise floor prior to encoding. This dithering intervention establishes a macro-variance floor that "smuggles" spatial geometry across the normalization layers. Evaluated on a ternary directional classification task (Left, Center, Right), we demonstrate that dithered models achieve exceptional spatial resolution--reaching up to 97.2% localization accuracy even on subtle 0.5 panning amplitudes--and demonstrates robust, zero-shot generalization to entirely unseen spatial configurations. Our results suggest that with the appropriate acoustic regularization, standard multimodal models are natively capable of generalized stereo audio understanding.

立体声感知大模型语音生成音频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。