arXiv:2511.15848cs.AIcs.CL2025-11被引 41

首个成功实现音频领域推理的模型,让语音理解更像人类思考。

Step-Audio-R1 Technical Report

  • 通过声学特征锚定的推理蒸馏框架,生成真实相关的音频推理链。
  • 性能超越Gemini 2.5 Pro,接近Gemini 3 Pro在多类音频任务上的表现。
  • 为多模态深度推理提供新路径,适合研究跨模态智能的学者。

近期推理模型在文本与视觉领域通过长链思维推理取得显著进展。然而,在音频语言模型中存在一个困惑现象:其表现往往在极少或无推理时更优,引发根本疑问——音频智能能否真正受益于深思熟虑?我们提出Step-Audio-R1,首个在音频领域成功解锁推理能力的模型。通过提出的模态锚定推理蒸馏(MGRD)框架,Step-Audio-R1学习生成真正基于声学特征的推理链,而非脱离音频事实的幻觉式思辨。该模型在涵盖语音、环境音和音乐的综合性音频理解与推理基准上表现出强大能力,超越Gemini 2.5 Pro,并达到与最先进模型Gemini 3 Pro相当的水平。结果表明,当合理锚定,推理能力可在模态间迁移,使长链思辨从负担变为音频智能的强大资产。Step-Audio-R1首次成功构建音频推理模型,为打造真正跨感官深度推理系统开辟新道路。

原文摘要 · Abstract (English)

Recent advances in reasoning models have demonstrated remarkable success in text and vision domains through extended chain-of-thought deliberation. However, a perplexing phenomenon persists in audio language models: they consistently perform better with minimal or no reasoning, raising a fundamental question - can audio intelligence truly benefit from deliberate thinking? We introduce Step-Audio-R1, the first audio reasoning model that successfully unlocks reasoning capabilities in the audio domain. Through our proposed Modality-Grounded Reasoning Distillation (MGRD) framework, Step-Audio-R1 learns to generate audio-relevant reasoning chains that genuinely ground themselves in acoustic features rather than hallucinating disconnected deliberations. Our model exhibits strong audio reasoning capabilities, surpassing Gemini 2.5 Pro and achieving performance comparable to the state-of-the-art Gemini 3 Pro across comprehensive audio understanding and reasoning benchmarks spanning speech, environmental sounds, and music. These results demonstrate that reasoning is a transferable capability across modalities when appropriately anchored, transforming extended deliberation from a liability into a powerful asset for audio intelligence. By establishing the first successful audio reasoning model, Step-Audio-R1 opens new pathways toward building truly multimodal reasoning systems that think deeply across all sensory modalities.

音频推理多模态推理蒸馏AI思考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。