通过定位语言差异层,用小规模适配提升阿拉伯语医疗大模型性能。
Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection
- 只在跨语言表征分化的中间层进行低秩适配,避免全网微调。
- 在多选题医学问答中超越全网LoRA和零样本基线,提升12.3个百分点。
- 适用于阿拉伯语医疗问答与对话,无需任务微调,适合资源匮乏语言研究者。
大型语言模型在英语医疗任务中表现优异,但在阿拉伯语中显著退化,通常归因于训练数据有限。我们通过调控探针和因果激活修补系统性检验这一假设,发现阿拉伯语医疗知识存在于模型中间层表示中,但未能在输出层显现。这一机制洞察推动了针对性适配策略:不微调整个网络,而是提出目标低秩适配(TLoRA),仅作用于跨语言表征开始分化的层窗口,位于输出层之前。我们在多项选择医学问答上评估TLoRA,结果优于全网LoRA、零样本和少样本基线。此外,在简答生成与多轮临床对话任务中,TLoRA也表现良好,无需任务特定微调。我们还构建了AraClinicDialog,一个由临床医生构建的现代标准阿拉伯语医疗对话基准,涵盖四种阿拉伯方言的验证版本。这些成果表明,机制诊断可为低资源语言医疗LLM的靶向适配提供实用指导。
原文摘要 · Abstract (English)
Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data. We systematically investigate this assumption via tuned lens probing and causal activation patching, and find that Arabic medical knowledge is present in intermediate model representations but fails to surface at the output. This mechanistic insight motivates a targeted adaptation strategy: rather than fine-tuning the full network, we propose Targeted Low-Rank Adaptation (TLoRA), restricted to the layer window where cross-lingual representations diverge, upstream of the output layers where the failure manifests. We evaluate TLoRA on multiple-choice medical QA, where our approach outperforms full-network LoRA, zero-shot, and few-shot baselines. We further evaluate it on short-answer generation and multi-turn clinical dialogue, where it performs competitively without the need for task-specific finetuning. We additionally introduce AraClinicDialog, a clinician-constructed Arabic medical dialogue benchmark in MSA with validated variants across four Arabic dialects. Together, these contributions demonstrate that mechanistic diagnosis can serve as a practical guide for targeted adaptation in underrepresented-language medical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。