构建中英文饮品点单测试集,评估语音助手真实场景下的理解能力。
StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario
- 收集真实语音语料,含停顿、更正等自然说话现象。
- 支持语音到槽位、转写到槽位等多种评估任务。
- 适合评估多语言语音助手在复杂场景下的鲁棒性。
大语言模型和语音助手在任务导向交互中应用日益广泛,但现有评估多依赖受控场景,难以反映真实用户请求的多样性和复杂性。以饮品点单为例,涉及多种命名实体、饮品类型、尺寸、个性化定制及品牌术语,还包含停顿、自我修正等自发性语音特征。为此,我们提出StarDrinks,一个包含英语和韩语语音语料、转写文本及标注槽位的测试集。该数据集支持语音到槽位(SLU)、转写到槽位(NLU)以及语音到转写(ASR)的评估,为模型在语言丰富、真实任务场景中的鲁棒性与泛化能力提供可靠基准。
原文摘要 · Abstract (English)
LLMs and speech assistants are increasingly used for task-oriented interactions, yet their evaluation often relies on controlled scenarios that fail to capture the variability and complexity of real user requests. Drink ordering, for example, involves diverse named entities, drink types, sizes, customizations, and brand-specific terminology, as well as spontaneous speech phenomena such as hesitations and self-corrections. To address this gap, we introduce StarDrinks, a test set in English and Korean containing speech utterances features, transcriptions, and annotated slots. Our dataset supports speech-to-slots SLU, transcription-to-slots NLU, and speech-to-transcription ASR evaluation, providing a realistic benchmark for model robustness and generalization in a linguistically rich, real-world task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。