arXiv:2605.19173cs.CL2026-05

不同提示语言影响大模型诊断表现,法语不如英语可靠。

Prompting language influences diagnostic reasoning and accuracy of large language models

论文配图:Prompting language influences diagnostic reasoning and accuracy of large language models
图 1 · 摘自论文原文
  • 用英法双语提示测试五款大模型的临床推理能力
  • 四款模型在英语下诊断准确率更高,差距达0.37至0.91分
  • 仅o3模型不受语言影响,适合多语言医疗场景

大型语言模型(LLMs)在临床决策支持中应用日益广泛,但多数评估局限于英语,其在其他语言中的可靠性尚不明确。本研究通过比较五款模型(o3、DeepSeek-R1、GPT-4-Turbo、Llama-3.1-405B-Instruct、BioMistral-7B)在英语与法语提示下的表现,评估了提示语言对诊断推理和最终诊断准确率的影响。共使用180个涵盖16个医学专科的临床案例,由两名医生依据18分量表评价诊断准确性与推理质量。结果显示,四款模型在英语提示下表现更优(均值差异0.37–0.91,校正后p < 0.05),且推理差异体现在鉴别诊断、逻辑结构和内部一致性等多个方面。o3是唯一未受语言影响的模型。研究证实,提示语言仍是决定大模型临床表现的关键因素,对全球范围内的公平语言文化部署具有重要意义。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly explored for clinical decision support, yet most evaluations are conducted in English, leaving their reliability in other languages uncertain. Here we evaluate the impact of prompting language on diagnostic reasoning and final diagnosis accuracy by comparing English and French performance across five LLMs (o3, DeepSeek-R1, GPT-4-Turbo, Llama-3.1-405B-Instruct, and BioMistral-7B). A total of 180 clinical vignettes covering 16 medical specialties were assessed by two physicians using an 18-point scale evaluating both diagnosis accuracy and reasoning quality. Four of the five models performed better in English (mean difference 0.37-0.91, adjusted p < 0.05), with the gap spanning multiple aspects of reasoning, including differential diagnosis, logical structure, and internal validity. o3 was the only model showing no overall language effect. These findings demonstrate that prompting language remains a critical determinant of LLM clinical performance, with implications for equitable linguistico-cultural deployment worldwide.

大模型临床诊断多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。