构建5000万条多语言语音指令数据集,提升语音大模型的指令理解能力。
SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
- 从公开语音语料生成5000万条语音-文本指令对,覆盖5种语言。
- 基于该数据集训练的SIFT-LLM在指令遵循任务上超越现有模型。
- 提供专门评估语音大模型指令能力的基准EvalSIFT,适合语音与NLP交叉研究者。
我们提出SIFT(Speech Instruction Fine-Tuning),一个包含5000万条样本的大型多语言语音指令微调数据集,用于语音-文本大语言模型(LLMs)的微调与预训练。SIFT-50M源自总计14,000小时公开语音语料,结合大语言模型及现成专家模型构建。数据集涵盖五种语言,涵盖多样化的语音理解与可控语音生成指令。基于SIFT-50M,我们训练了SIFT-LLM,在指令遵循基准测试中表现优于现有语音-文本大模型,同时在基础语音任务上保持竞争力。为支持进一步研究,我们还推出了EvalSIFT,一个专为评估语音-文本大模型指令遵循能力而设计的基准数据集。
原文摘要 · Abstract (English)
We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50M is built from publicly available speech corpora, which collectively contain 14K hours of speech, and leverages LLMs along with off-the-shelf expert models. The dataset spans five languages, encompassing a diverse range of speech understanding as well as controllable speech generation instructions. Using SIFT-50M, we train SIFT-LLM, which outperforms existing speech-text LLMs on instruction-following benchmarks while achieving competitive performance on foundational speech tasks. To support further research, we also introduce EvalSIFT, a benchmark dataset specifically designed to evaluate the instruction-following capabilities of speech-text LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。