arXiv:2601.21124cs.SDcs.AI2026-01被引 4

让大模型理解任意麦克风阵列的声源位置,突破设备限制。

PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs

  • 用原始多通道音频和麦克风坐标直接生成空间嵌入,不依赖固定布局。
  • 在跨设备场景下定位准确率领先现有方法,支持任意麦克风阵列。
  • 首次实现大模型基于空间音频进行复杂推理与定向语音转写。

当前多模态大模型将音频处理为单声道流,忽略了具身智能所需的丰富空间信息。现有空间音频模型则受限于固定麦克风布局,难以适配不同设备。我们提出PhaseCoder,一种仅基于Transformer的空间音频编码器,对麦克风几何结构无感。该模型输入原始多通道音频及麦克风坐标,完成声源定位并生成鲁棒的空间嵌入。实验表明,可对Gemma 3n大模型进行微调,使其理解由PhaseCoder生成的“空间音频标记”。我们在麦克风无关的定位基准上达到当前最优性能,并首次实现大模型从任意麦克风阵列执行复杂空间推理与定向语音转写任务。

原文摘要 · Abstract (English)

Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained to fixed microphone geometries, preventing deployment across diverse devices. We present PhaseCoder, a transformer-only spatial audio encoder that is agnostic to microphone geometry. PhaseCoder takes raw multichannel audio and microphone coordinates as inputs to perform localization and produces robust spatial embeddings. We demonstrate that Gemma 3n LLM can be fine-tuned to reason over "Spatial Audio Tokens" produced by PhaseCoder. We show our encoder achieves state-of-the-art results on microphone-invariant localization benchmarks and, for the first time, enables an LLM to perform complex spatial reasoning and targeted transcription tasks from an arbitrary microphone array.

空间音频多模态大模型声源定位麦克风阵列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。