通过双语协同推理,提升低资源语言医学问答准确性
Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

- 让英文与本地语言并行推理,用概念对齐融合临床知识
- 在7种语言上平均提升5%性能,低资源语言增益更明显
- 适合医疗AI多语言落地,尤其关注本土化临床实践的团队
尽管增强推理的大语言模型在英语医学任务中表现优异,但多语言差距依然存在,本地语言的推理能力显著较弱,限制了全球医疗部署的公平性。为弥合这一差距,我们提出Med-CoReasoner,一种语言感知的协同推理框架:同时激发英文与本地语言的推理路径,将二者抽象为结构化概念,并通过概念级对齐与检索,将本地临床知识融入英文逻辑框架。该设计结合了英文推理的结构稳健性与本地语言中的实践型专业知识。为评估超越多项选择题的多语言医学推理能力,我们构建了MultiMed-X基准,涵盖七种语言,包含专家标注的长文本问答与自然语言推理任务,每语言350个实例。三组基准实验显示,Med-CoReasoner在多语言推理上平均提升5%,尤其在低资源语言中表现突出。模型蒸馏与专家评估进一步证实,Med-CoReasoner生成的推理链条具有临床合理性与文化适配性。
原文摘要 · Abstract (English)
While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local languages, limiting equitable global medical deployment. To bridge this gap, we introduce Med-CoReasoner, a language-informed co-reasoning framework that elicits parallel English and local-language reasoning, abstracts them into structured concepts, and integrates local clinical knowledge into an English logical scaffold via concept-level alignment and retrieval. This design combines the structural robustness of English reasoning with the practice-grounded expertise encoded in local languages. To evaluate multilingual medical reasoning beyond multiple-choice settings, we construct MultiMed-X, a benchmark covering seven languages with expert-annotated long-form question answering and natural language inference tasks, comprising 350 instances per language. Experiments across three benchmarks show that Med-CoReasoner improves multilingual reasoning performance by an average of 5%, with particularly substantial gains in low-resource languages. Moreover, model distillation and expert evaluation analysis further confirm that Med-CoReasoner produces clinically sound and culturally grounded reasoning traces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。