测试大模型在多语言下回答健康问题的一致性,发现存在严重信息偏差。
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
- 构建多语言健康问答数据集,涵盖四种语言并标注疾病类别
- 发现不同语言间回答差异显著,可能传播医疗错误信息
- 提出新评估流程,可精准比较双语间回答一致性
公平获取可靠健康信息对公共健康至关重要,但在线健康资源的质量因语言而异,引发人们对大型语言模型在医疗领域表现一致性的担忧。本研究考察了大模型在英语、德语、土耳其语和中文下对健康相关问题的回答一致性。我们大幅扩展了HealthFC数据集,按疾病类型分类健康问题,并新增土耳其语和中文翻译以拓展多语言覆盖范围。研究揭示了回答中存在显著不一致,可能加剧医疗信息误导。主要贡献包括:1)带有疾病类别元信息的多语言健康咨询数据集;2)一种基于提示的新评估工作流,可通过解析实现两种语言间的子维度对比。研究结果凸显了在多语言场景部署基于大模型工具的关键挑战,强调需加强跨语言对齐以保障医疗信息的准确性与公平性。
原文摘要 · Abstract (English)
Equitable access to reliable health information is vital for public health, but the quality of online health resources varies by language, raising concerns about inconsistencies in Large Language Models (LLMs) for healthcare. In this study, we examine the consistency of responses provided by LLMs to health-related questions across English, German, Turkish, and Chinese. We largely expand the HealthFC dataset by categorizing health-related questions by disease type and broadening its multilingual scope with Turkish and Chinese translations. We reveal significant inconsistencies in responses that could spread healthcare misinformation. Our main contributions are 1) a multilingual health-related inquiry dataset with meta-information on disease categories, and 2) a novel prompt-based evaluation workflow that enables sub-dimensional comparisons between two languages through parsing. Our findings highlight key challenges in deploying LLM-based tools in multilingual contexts and emphasize the need for improved cross-lingual alignment to ensure accurate and equitable healthcare information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。