arXiv:2510.05542cs.SDcs.CL2025-10被引 4

首个能完整描述声音场景的空间音频大模型,支持多源定位与环境还原。

Sci-Phi: A Large Language Model Spatial Audio Descriptor

  • 双编码器架构融合空间与频谱信息,实现多声源同步解析。
  • 单次处理可识别最多4个方向声源及背景环境参数,准确率超基准30%以上。
  • 实测在真实混响环境下表现稳定,适合智能语音、虚拟现实等场景。

声景感知需描述声音类型、时间、方向、距离、音量及混响特性。尽管音频语言模型在声音识别上表现优异,但单通道输入严重限制了空间理解能力。本文提出Sci-Phi,一种具备双空间与频谱编码器的音频大语言模型,可估计所有声源及环境的完整参数。模型基于超过4000小时合成的一阶Ambisonics音频数据(含元数据)训练,一次处理可枚举并描述最多四个方向性声源,同时识别非方向性背景音与房间特性。采用置换不变评估协议,使用15项指标覆盖内容、位置、时间、音量与混响。实验分析了声源数量、信噪比、混响水平及相似声源混合下的鲁棒性。结果表明,Sci-Phi在真实房间冲击响应下仅出现轻微性能下降。本工作首次实现音频大模型对完整声景的描述,具备实际部署潜力。

原文摘要 · Abstract (English)

Acoustic scene perception involves describing the type of sounds, their timing, their direction and distance, as well as their loudness and reverberation. While audio language models excel in sound recognition, single-channel input fundamentally limits spatial understanding. This work presents Sci-Phi, a spatial audio large language model with dual spatial and spectral encoders that estimates a complete parameter set for all sound sources and the surrounding environment. Learning from over 4,000 hours of synthetic first-order Ambisonics recordings including metadata, Sci-Phi enumerates and describes up to four directional sound sources in one pass, alongside non-directional background sounds and room characteristics. We evaluate the model with a permutation-invariant protocol and 15 metrics covering content, location, timing, loudness, and reverberation, and analyze its robustness across source counts, signal-to-noise ratios, reverberation levels, and challenging mixtures of acoustically, spatially, or temporally similar sources. Notably, Sci-Phi generalizes to real room impulse responses with only minor performance degradation. Overall, this work establishes the first audio LLM capable of full spatial-scene description, with strong potential for real-world deployment. Demo: https://sci-phi-audio.github.io/demo

音频生成空间音频大模型声源分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。