arXiv:2512.23808cs.CLcs.SD2025-12被引 117

超大规模音频预训练让模型仅用少量样本就能完成新任务

MiMo-Audio: Audio Language Models are Few-Shot Learners

  • 用超过1亿小时数据训练,实现无需微调的少样本学习
  • 在语音智能与音频理解任务中达到开源模型最好水平
  • 能处理语音转换、风格迁移等未训练过的新型任务

现有音频语言模型通常依赖特定任务微调,而人类仅需少量示例或简单指令即可泛化。GPT-3表明,扩大文本下一词预测预训练可带来强泛化能力,我们认为该范式同样适用于音频领域。通过将MiMo-Audio的预训练数据扩展至超过一亿小时,我们观察到多种音频任务中少样本学习能力的涌现。我们系统评估了这些能力,发现MiMo-Audio-7B-Base在开源模型中于语音智能与音频理解基准上达到最佳表现。除标准指标外,该模型还能泛化至训练数据中不存在的任务,如语音转换、风格迁移和语音编辑。其语音续写能力强大,可生成高度真实的脱口秀、朗诵、直播及辩论场景。在后训练阶段,我们构建多样化指令调优语料,并在音频理解和生成中引入思考机制。MiMo-Audio-7B-Instruct在音频理解(MMSU, MMAU, MMAR, MMAU-Pro)、口语对话(Big Bench Audio, MultiChallenge Audio)及指令式语音合成评测中达到开源模型最优,接近甚至超越闭源模型。模型检查点与完整评估套件见https://github.com/XiaomiMiMo/MiMo-Audio。

原文摘要 · Abstract (English)

Existing audio language models typically rely on task-specific fine-tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with only a few examples or simple instructions. GPT-3 has shown that scaling next-token prediction pretraining enables strong generalization capabilities in text, and we believe this paradigm is equally applicable to the audio domain. By scaling MiMo-Audio's pretraining data to over one hundred million of hours, we observe the emergence of few-shot learning capabilities across a diverse set of audio tasks. We develop a systematic evaluation of these capabilities and find that MiMo-Audio-7B-Base achieves SOTA performance on both speech intelligence and audio understanding benchmarks among open-source models. Beyond standard metrics, MiMo-Audio-7B-Base generalizes to tasks absent from its training data, such as voice conversion, style transfer, and speech editing. MiMo-Audio-7B-Base also demonstrates powerful speech continuation capabilities, capable of generating highly realistic talk shows, recitations, livestreaming and debates. At the post-training stage, we curate a diverse instruction-tuning corpus and introduce thinking mechanisms into both audio understanding and generation. MiMo-Audio-7B-Instruct achieves open-source SOTA on audio understanding benchmarks (MMSU, MMAU, MMAR, MMAU-Pro), spoken dialogue benchmarks (Big Bench Audio, MultiChallenge Audio) and instruct-TTS evaluations, approaching or surpassing closed-source models. Model checkpoints and full evaluation suite are available at https://github.com/XiaomiMiMo/MiMo-Audio.

音频大模型少样本学习语音生成指令调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。