构建多语言语音指令数据集,真实评估语音大模型的指令遵循能力。
Do What I Say: A Spoken Prompt Dataset for Instruction-Following
- 设计9个任务、11种语言的语音与文字并列指令数据集。
- 语音指令在低资源和跨语言场景下表现显著逊于文本指令。
- 仅在输出为语音的任务中,语音指令表现接近文本指令。
语音大语言模型(SLLMs)发展迅速,但现有评估多基于文本提示,难以反映真实语音交互场景。为此,我们提出DoWhatISay(DOWIS)数据集,包含9项任务、11种语言的人类录制语音与文字指令,每对任务-语言组合提供10种变体及5种风格。利用该数据集,我们对主流SLLMs进行评估,分析提示模态、风格、语言与任务类型之间的相互影响。结果表明,文本提示在各类场景中均优于语音提示,尤其在低资源和跨语言设置下差距明显;仅在语音输出任务中,语音提示性能接近文本提示,凸显了在真实场景中使用语音指令评估SLLMs的必要性。
原文摘要 · Abstract (English)
Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompts, which may not reflect real-world scenarios where users interact with speech. To address this gap, we introduce DoWhatISay (DOWIS), a multilingual dataset of human-recorded spoken and written prompts designed to pair with any existing benchmark for realistic evaluation of SLLMs under spoken instruction conditions. Spanning 9 tasks and 11 languages, it provides 10 prompt variants per task-language pair, across five styles. Using DOWIS, we benchmark state-of-the-art SLLMs, analyzing the interplay between prompt modality, style, language, and task type. Results show that text prompts consistently outperform spoken prompts, particularly for low-resource and cross-lingual settings. Only for tasks with speech output, spoken prompts do close the gap, highlighting the need for speech-based prompting in SLLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。