首个多语言口语问答数据集,评估大模型真实对话能力。
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
- 构建3.3万条多语种自然口语问答,覆盖低资源语言和方言。
- 首次在真实语音交互场景下评估大模型表现,揭示语音差异影响。
- 适合研究多语言语音理解、大模型评测的学者使用。
大型语言模型(LLMs)在多个领域表现出色,但针对多语言口语查询的评测仍属空白。本文提出SpokenNativQA,首个面向真实对话场景的多语言、文化对齐口语问答数据集,包含约3.3万条自然口语问答,涵盖多种语言,包括低资源语言与方言,有效弥补文本问答数据集在语音变异性、口音和语言多样性方面的不足。我们评估了多种ASR系统与LLMs在口语问答任务中的表现,并公开数据集(https://huggingface.co/datasets/QCRI/SpokenNativQA)及实验脚本(https://llmebench.qcri.org/),供研究社区使用。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable performance across various disciplines and tasks. However, benchmarking their capabilities with multilingual spoken queries remains largely unexplored. In this study, we introduce SpokenNativQA, the first multilingual and culturally aligned spoken question-answering (SQA) dataset designed to evaluate LLMs in real-world conversational settings. The dataset comprises approximately 33,000 naturally spoken questions and answers in multiple languages, including low-resource and dialect-rich languages, providing a robust benchmark for assessing LLM performance in speech-based interactions. SpokenNativQA addresses the limitations of text-based QA datasets by incorporating speech variability, accents, and linguistic diversity. We benchmark different ASR systems and LLMs for SQA and present our findings. We released the data at (https://huggingface.co/datasets/QCRI/SpokenNativQA) and the experimental scripts at (https://llmebench.qcri.org/) for the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。