arXiv:2609.07687cs.CL2026-09

医学大模型该跨语言一致回答,还是按文化调整?

Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions

论文配图:Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions
图 1 · 摘自论文原文
  • 从一致性与文化适应双视角梳理多语言医疗NLP研究
  • 356名专家调查显示医界对跨语言一致性意见分歧
  • 现有模型无法模拟真实人群的文化偏好差异

多语言大模型在回答医疗问题时,应保持跨语言答案一致,还是根据文化背景调整?现有医疗多语言基准通常假设医学正确答案应跨语言一致,将语言差异视为模型错误。但文化适应研究认为,恰当的医疗建议可能因语境而异。本文通过一致性与文化适应两个视角综述多语言医疗NLP文献,识别出三大空白:缺乏利益相关方(如医疗、NLP、人类学)视角,缺乏实证证据证明哪种方式更利于用户,以及缺少能区分普遍正确与文化特异性案例的基准。为填补第一项空白,我们在德国、西班牙和美国对356名医疗、NLP和人类学专业人士进行调查。人类学家普遍支持文化适应,而医疗与NLP受访者意见分裂,尤其美欧医疗人员存在显著差异。当用专业身份与国家角色提示大模型时,其无法复现这种多样性,高估了跨语言一致性偏好。结论是:目前尚无明确证据表明一致性或适应性更优,亟需实证研究验证不同文化背景下哪种方式更有利于用户。

原文摘要 · Abstract (English)

Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume that medically correct answers should remain consistent across languages and treat cross-lingual variation as model error. In contrast, cultural adaptation research argues that appropriate medical answers may legitimately differ across contexts. We review the multilingual medical NLP literature through these two perspectives, we identify three gaps: limited stakeholder perspectives (e.g., of medical professionals), a lack of empirical evidence on which approach better serves users, and no benchmarks capable of distinguishing universally correct from culture-specific cases. To address the first gap, we survey 356 participants across three stakeholder groups (medical, NLP, and anthropology professionals) in three countries (Germany, Spain, and the United States). Anthropologists consistently favor adaptation, while medical and NLP respondents remain divided, with notable divergence between U.S. and European medical professionals. LLMs prompted with profession and country personas fail to reproduce this variation, overestimating cross-lingual consistency preference among NLP and medical personas. We conclude that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.

多语言医疗AI文化适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。