小型多模态语音对话模型,本地运行且性能领先。
Voxtral
- 融合语音与文本的多模态训练,支持长音频理解
- 小模型超越多个闭源产品,32K上下文可处理40分钟音频
- 适合需要本地部署的语音助手开发者
我们提出Voxtral Mini和Voxtral Small两款多模态语音对话模型。Voxtral具备理解语音与文本文档的能力,在多样化的音频基准测试中达到当前最优表现,同时保持强大的文本处理能力。Voxtral Small在性能上超越多个闭源模型,且体积足够小可在本地运行。32K上下文窗口支持长达40分钟的音频文件及长轮次多轮对话。我们还贡献了三个用于评估语音理解模型在知识与趣味问答任务上的新基准。两款模型均采用Apache 2.0许可证开源。
原文摘要 · Abstract (English)
We present Voxtral Mini and Voxtral Small, two multimodal audio chat models. Voxtral is trained to comprehend both spoken audio and text documents, achieving state-of-the-art performance across a diverse range of audio benchmarks, while preserving strong text capabilities. Voxtral Small outperforms a number of closed-source models, while being small enough to run locally. A 32K context window enables the model to handle audio files up to 40 minutes in duration and long multi-turn conversations. We also contribute three benchmarks for evaluating speech understanding models on knowledge and trivia. Both Voxtral models are released under Apache 2.0 license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。