arXiv:2601.09385cs.SDcs.CL2026-01被引 11

开源多模态大模型框架SLAM-LLM,专注语音音频音乐处理。

SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing

  • 模块化设计支持语音、音频、音乐等多种模态输入
  • 提供语音识别、音频描述等任务的高性能预训练模型
  • 适合从事语音与音乐方向研究的开发者快速上手

近期开源的多模态大语言模型(MLLM)框架(如LLaVA)为开发者提供了便利。但多数框架以视觉为主,对语音、音频和音乐的支持有限,导致研究人员需耗费大量精力进行代码编写与超参数调优。我们提出SLAM-LLM,一个专注于语音、语言、音频和音乐处理的开源深度学习框架。该框架采用模块化配置,集成多种编码器、投影器、大语言模型及参数高效微调插件,并提供主流任务的详细训练与推理方案。包含基于LLM的自动语音识别(ASR)、自动化音频描述(AAC)和音乐描述(MC)的高性能检查点。部分方案已达到或接近当前最优性能,相关技术已被学术论文接收。我们希望借助此开源框架加速音频类多模态大模型的研究迭代与开发进程,呼吁社区共同推进语音、音频与音乐的LLM应用。

原文摘要 · Abstract (English)

The recent surge in open-source Multimodal Large Language Models (MLLM) frameworks, such as LLaVA, provides a convenient kickoff for artificial intelligence developers and researchers. However, most of the MLLM frameworks take vision as the main input modality, and provide limited in-depth support for the modality of speech, audio, and music. This situation hinders the development of audio-language models, and forces researchers to spend a lot of effort on code writing and hyperparameter tuning. We present SLAM-LLM, an open-source deep learning framework designed to train customized MLLMs, focused on speech, language, audio, and music processing. SLAM-LLM provides a modular configuration of different encoders, projectors, LLMs, and parameter-efficient fine-tuning plugins. SLAM-LLM also includes detailed training and inference recipes for mainstream tasks, along with high-performance checkpoints like LLM-based Automatic Speech Recognition (ASR), Automated Audio Captioning (AAC), and Music Captioning (MC). Some of these recipes have already reached or are nearing state-of-the-art performance, and some relevant techniques have also been accepted by academic papers. We hope SLAM-LLM will accelerate iteration, development, data engineering, and model training for researchers. We are committed to continually pushing forward audio-based MLLMs through this open-source framework, and call on the community to contribute to the LLM-based speech, audio and music processing.

多模态语音处理开源框架音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。