构建多语言医学推理数据,提升大模型跨语言医疗问答能力
Multilingual Medical Reasoning for Question Answering with Large Language Models
- 从维基百科提取医学知识,生成英意西三语推理链
- 在80亿参数模型上实现多语言医疗问答新纪录
- 适合开发透明可靠的跨国临床辅助系统
具备推理能力的大语言模型在医疗问答任务中展现出巨大潜力。现有方法主要聚焦英文,依赖通用大模型蒸馏,存在医学知识可靠性问题。本文基于维基百科医学信息,采用检索增强生成方法,构建了50万条英、意、西三语医学推理链,用于解答来自MedQA和MedMCQA的医学问题,并将这两个数据集扩展至意大利语和西班牙语。我们在域内与域外设置下测试该流程,在多个医疗问答基准上均表现优异,无论是上下文学习(少样本)还是监督微调均有效提升性能,达到80亿参数模型的最先进水平。我们公开发布全部资源:推理链、翻译后的问答数据集、Medical-Wikipedia及微调模型,以支持多语言医疗决策系统的透明化发展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) with reasoning capabilities have recently demonstrated strong potential in medical Question Answering (QA). Existing approaches are largely English-focused and primarily rely on distillation from general-purpose LLMs, raising concerns about the reliability of their medical knowledge. In this work, we present a method to generate multilingual reasoning traces based on medical knowledge extracted from Wikipedia. We produce 500k traces in English, Italian, and Spanish, using a retrieval-augmented generation approach over medical information from Wikipedia. The traces are generated to solve medical questions drawn from MedQA and MedMCQA, which we extend to Italian and Spanish. We test our pipeline in both in-domain and out-of-domain settings across Medical QA benchmarks, and demonstrate that our reasoning traces improve performance both when utilized via in-context learning (few-shot) and supervised fine-tuning, yielding state-of-the-art results among 8B-parameter LLMs. We believe that these resources can support the development of more transparent clinical decision-support tools in multilingual settings. We release the full suite of resources: reasoning traces, translated QA datasets, Medical-Wikipedia, and fine-tuned models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。