arXiv:2412.14304cs.CLcs.AI2024-12中稿 · AAAI被引 15

首个多语言眼科问答基准,发现大模型存在语言偏见并提出新去偏方法。

Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs

  • 构建跨语言眼科问答数据集,支持多语种直接对比。
  • 6个主流大模型在7种语言中表现差异显著,存在明显语言偏见。
  • 提出CLARA系统,通过检索增强与自我验证减少多语言偏差,适合医疗公平性研究者。

当前眼科临床流程面临过度转诊、等待时间长及病历复杂异质等问题。大型语言模型(LLMs)有望自动化分诊、视力评估和报告摘要等环节。然而,现有LLMs在自然语言问答任务中跨语言表现差异巨大,可能加剧低收入和中等收入国家(LMICs)的医疗不平等。本研究首次构建了多语言眼科问答基准,包含人工标注的平行问题,支持跨语言直接比较。对7种语言下6个主流LLMs的评估显示,各语言间性能差异显著,存在严重语言偏见,提示其在LMICs临床部署的风险。现有去偏方法如翻译链思维或检索增强生成(RAG)单独使用时难以弥合差距,且缺乏医学领域针对性。为此,我们提出CLARA(跨语言反思代理系统),一种基于检索增强生成与自我验证的新型推理阶段去偏方法。该方法不仅提升所有语言下的性能,还显著缩小多语言偏见差距,促进全球范围内大模型的公平应用。

原文摘要 · Abstract (English)

Current ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language models (LLMs) present a promising solution to automate various procedures such as triaging, preliminary tests like visual acuity assessment, and report summaries. However, LLMs have demonstrated significantly varied performance across different languages in natural language question-answering tasks, potentially exacerbating healthcare disparities in Low and Middle-Income Countries (LMICs). This study introduces the first multilingual ophthalmological question-answering benchmark with manually curated questions parallel across languages, allowing for direct cross-lingual comparisons. Our evaluation of 6 popular LLMs across 7 different languages reveals substantial bias across different languages, highlighting risks for clinical deployment of LLMs in LMICs. Existing debiasing methods such as Translation Chain-of-Thought or Retrieval-augmented generation (RAG) by themselves fall short of closing this performance gap, often failing to improve performance across all languages and lacking specificity for the medical domain. To address this issue, We propose CLARA (Cross-Lingual Reflective Agentic system), a novel inference time de-biasing method leveraging retrieval augmented generation and self-verification. Our approach not only improves performance across all languages but also significantly reduces the multilingual bias gap, facilitating equitable LLM application across the globe.

多语言眼科大模型去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。