arXiv:2603.09556cs.CL2026-03被引 3

让语音模型更懂推理,生成自然回应。

ALARM: Audio-Language Alignment for Reasoning Models

  • 自重述机制将文本回复转为适配语音理解的版本。
  • 用600万条数据训练,40亿参数模型超越多数更大模型。
  • 适合需要语音推理能力的研究者和开发者。

大音频语言模型(ALMs)通过引入听觉理解扩展了大型语言模型(LLMs)。常见方法是冻结LLM,仅训练适配器处理自生成目标。然而,对于具有链式思维(CoT)能力的推理型语言模型(RLMs),其内部推理过程会暴露文本代理输入,导致响应不自然。本文提出自重述机制,将自生成的回答转化为适配语音理解的版本,同时保持分布对齐。进一步融合并压缩多个音频编码器以获得更强表征。训练中构建了一个包含600万实例的多任务语料库(250万唯一提示),覆盖19,000小时的语音、音乐与音效。我们提出的40亿参数ALM在同类规模模型中表现领先,并在多个音频推理基准上超越多数更大模型,同时保持文本能力,训练成本低。尤其在MMAU-speech和MMSU基准上取得开源最佳结果,位列所有模型第三。

原文摘要 · Abstract (English)

Large audio language models (ALMs) extend LLMs with auditory understanding. A common approach freezes the LLM and trains only an adapter on self-generated targets. However, this fails for reasoning LLMs (RLMs) whose built-in chain-of-thought traces expose the textual surrogate input, yielding unnatural responses. We propose self-rephrasing, converting self-generated responses into audio-understanding variants compatible with RLMs while preserving distributional alignment. We further fuse and compress multiple audio encoders for stronger representations. For training, we construct a 6M-instance multi-task corpus (2.5M unique prompts) spanning 19K hours of speech, music, and sound. Our 4B-parameter ALM outperforms similarly sized models and surpasses most larger ALMs on related audio-reasoning benchmarks, while preserving textual capabilities with a low training cost. Notably, we achieve the best open-source result on the MMAU-speech and MMSU benchmarks and rank third among all the models.

语音理解推理模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。