构建多语言医疗大模型可信度评估基准,揭示其在事实性、偏见和隐私上的缺陷。
CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
- 设计涵盖15种语言的多维度评估框架,覆盖医疗问答、诊断、用药等18项任务。
- 发现主流大模型在低资源语言中事实正确率下降,且存在显著群体偏见与隐私漏洞。
- 适合关注全球医疗AI公平性与安全性的研究者与开发者参考。
将语言模型(LMs)融入医疗系统有望提升医疗流程与决策效率,但其可信度缺乏可靠评估,尤其在多语言医疗场景下尤为突出。现有语言模型主要在高资源语言上训练,难以应对中低资源语言中医疗查询的复杂性与多样性,限制了其在全球医疗中的部署。为此,本文提出CLINIC——一个综合性多语言医疗可信度评估基准。该基准系统评估语言模型在真实性、公平性、安全性、鲁棒性和隐私保护五个关键维度的表现,涵盖18项多样化任务,覆盖15种语言(遍及各大洲),涉及疾病状况、预防措施、诊断测试、治疗方案、手术及药物等广泛医疗主题。大量实验证明,当前语言模型在事实准确性方面表现不佳,对不同人口和语言群体存在偏见,并易受隐私泄露与对抗攻击影响。该研究揭示了现有模型的严重短板,为提升医疗语言模型在多元语言环境下的可信赖性与安全性提供了基础支撑。
原文摘要 · Abstract (English)
Integrating language models (LMs) in healthcare systems holds great promise for improving medical workflows and decision-making. However, a critical barrier to their real-world adoption is the lack of reliable evaluation of their trustworthiness, especially in multilingual healthcare settings. Existing LMs are predominantly trained in high-resource languages, making them ill-equipped to handle the complexity and diversity of healthcare queries in mid- and low-resource languages, posing significant challenges for deploying them in global healthcare contexts where linguistic diversity is key. In this work, we present CLINIC, a Comprehensive Multilingual Benchmark to evaluate the trustworthiness of language models in healthcare. CLINIC systematically benchmarks LMs across five key dimensions of trustworthiness: truthfulness, fairness, safety, robustness, and privacy, operationalized through 18 diverse tasks, spanning 15 languages (covering all the major continents), and encompassing a wide array of critical healthcare topics like disease conditions, preventive actions, diagnostic tests, treatments, surgeries, and medications. Our extensive evaluation reveals that LMs struggle with factual correctness, demonstrate bias across demographic and linguistic groups, and are susceptible to privacy breaches and adversarial attacks. By highlighting these shortcomings, CLINIC lays the foundation for enhancing the global reach and safety of LMs in healthcare across diverse languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。