arXiv:2510.17476cs.CL2025-10中稿 · NLDB 2026: The pap…被引 2

多语言医疗问答中,英文知识占优,影响AI可靠性。

Zoom In Disparities in Healthcare LLM Q&A

  • 构建多语言医疗数据集MultiWikiHealthCare,覆盖英德土汉意五语
  • 不同语言下大模型回答与维基百科对齐度差异显著,英语最优
  • 引入非英语上下文可有效提升非英语回答的准确性,适合跨语言医疗研究

公平获取可靠健康信息在人工智能融入医疗领域至关重要。然而,多语言大模型在信息质量上存在差异,引发对可靠性与一致性担忧。本文系统考察了英语、德语、土耳其语、中文(普通话)和意大利语在多语言医疗问答中预训练数据来源及事实一致性方面的跨语言差异。我们(一)构建了多语言医疗维基百科数据集MultiWikiHealthCare;(二)分析跨语言医疗内容覆盖率;(三)评估大模型回答与参考文献的一致性;(四)通过上下文信息与检索增强生成(RAG)开展事实对齐案例研究。结果表明,维基百科内容覆盖与大模型事实对齐均存在显著跨语言差异。所有大模型的回答更倾向于匹配英文维基百科,即使提示为非英语。在推理时提供非英语维基百科上下文片段,能有效引导事实对齐至文化相关知识。这些发现为构建更公平的多语言医疗AI系统提供了可行路径。

原文摘要 · Abstract (English)

Equitable access to reliable health information is vital when integrating AI into healthcare. Yet, information quality varies across languages, raising concerns about the reliability and consistency of multilingual Large Language Models (LLMs). We systematically examine cross-lingual disparities in pre-training source and factuality alignment in LLM answers for multilingual healthcare Q&A across English, German, Turkish, Chinese (Mandarin), and Italian. We (i) constructed Multilingual Wiki Health Care (MultiWikiHealthCare), a multilingual dataset from Wikipedia; (ii) analyzed cross-lingual healthcare coverage; (iii) assessed LLM response alignment with these references; and (iv) conducted a case study on factual alignment through the use of contextual information and Retrieval-Augmented Generation (RAG). Our findings reveal substantial cross-lingual disparities in both Wikipedia coverage and LLM factual alignment. Across LLMs, responses align more with English Wikipedia, even when the prompts are non-English. Providing contextual excerpts from non-English Wikipedia at inference time effectively shifts factual alignment toward culturally relevant knowledge. These results highlight practical pathways for building more equitable, multilingual AI systems for healthcare.

多语言医疗AI事实对齐RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。