arXiv:2507.16724cs.SDeess.AS2025-07被引 8

让语音模型理解声音方向并用文字编辑,突破传统音频语言模型局限

SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing

  • 用结构化嵌入拆分声音的语义与空间特征,实现多模态对齐
  • 零样本识别声音方向,支持基于文本的定向音频编辑
  • 适合做智能音频编辑、虚拟现实语音交互的研究者

空间音频理解对准确感知和解析声学环境至关重要。然而,现有音频-语言模型在处理空间音频和感知空间声景方面存在局限。为此,我们提出空间音频语言模型(SALM),通过多模态对比学习连接空间音频与自然语言。SALM融合文本编码器与双分支音频编码器,利用结构化音频嵌入将空间声音分解为语义与空间成分。其关键特性包括:空间音频与自然语言的无缝对齐、空间与语义表征的分离与联合提取、零样本方向分类,以及灵活的空间音频编辑支持。实验表明,SALM能有效捕捉并对齐跨模态表示,生成结构良好的音频嵌入。此外,SALM实现了高级编辑能力,如通过文本嵌入修改定向音频。

原文摘要 · Abstract (English)

Spatial audio understanding is essential for accurately perceiving and interpreting acoustic environments. However, existing audio-language models exhibit limitations in processing spatial audio and perceiving spatial acoustic scenes. To address this gap, we propose the Spatial Audio Language Model (SALM), a novel framework that bridges spatial audio and language through multi-modal contrastive learning. SALM integrates a text encoder with a dual-branch audio encoder that decomposes spatial sound into semantic and spatial components via structured audio embeddings. Key features of SALM include seamless alignment between spatial audio and natural language, both separate and joint extraction of spatial and semantic representations, zero-shot direction classification, and flexible support for spatial audio editing. Experimental results demonstrate that SALM effectively captures and aligns cross-modal representations, yielding well-structured audio embeddings. Furthermore, SALM enables advanced editing capabilities, such as modifying directional audio using text-based embeddings.

空间音频音频编辑多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。