构建多语言语音对话数据集,助力医疗信息查询系统研发
Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

- 基于世卫组织内容构建6000条多语言对话,含真实用户语音
- 覆盖阿拉伯/中文/英语/西班牙语,总计163小时语音数据
- 适合研究多语言对话系统、医疗问答与公平性评估的学者
构建语音对话数据集方法复杂,尤其在大规模多语言、多平行场景下更具挑战。本文提出HEALTHDIAL,一个大规模、多语言、多平行的数据集,用于开发与评估基于检索增强生成(RAG)的语音对话系统。该数据集包含6000条信息查询对话(每种语言1500条),内容基于世界卫生组织(WHO)的可信资料,涵盖阿拉伯语、汉语、英语和西班牙语四种官方语言,共收录163小时来自不同方言母语者的用户语音。每位说话者均标注了人口统计学(如性别、年龄)与社会语言学(如主要语言、原籍地区)信息。我们报告了关键对话任务的基准结果,发现即使在高资源语言间也存在持续性能差异。为支持后续研究,我们公开数据集、原型系统及数据采集与评估工具包。
原文摘要 · Abstract (English)
Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-parallel datasets at scale. This work introduces HEALTHDIAL, a large-scale, multilingual, and multi-parallel dataset for developing and evaluating retrieval-augmented generation (RAG)-based spoken dialogue systems. The dataset comprises 6,000 information-seeking dialogues (1,500 per language) grounded in trusted content from the World Health Organization (WHO) and 163 hours of user speech recorded from native speakers of diverse dialects across four official WHO languages: Arabic, Chinese, English, and Spanish. Each speaker is annotated with demographic (e.g., gender, age) and sociolinguistic (e.g., primary language, region of origin) variables. We report benchmark results across key dialogue tasks, which reveal consistent performance disparities across languages, even among high-resource ones. To support future research, we release the dataset, a prototype system, and a toolkit for data collection and system evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。