arXiv:2606.19266cs.CLcs.AI2026-06ACL

对比法语医学问答中不同模型适配方法,给出实用选择建议。

Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA

论文配图:Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA
图 1 · 摘自论文原文
  • 用持续预训练、微调及其组合测试三种模型家族和尺寸。
  • 法语医学多选题中微调表现最佳,开放问答则持续预训练更优。
  • 适合资源有限时做医疗大模型本地化,尤其关注法语场景。

大型语言模型的发展促使人们关注其在专业领域和语言上的适配,但领域适配策略的有效性仍不明确。本文以法语医学问答(QA)为案例,比较持续预训练(CPT)、监督微调(SFT)及其组合在三种模型族、多种尺寸和三种初始化方式下的表现,明确分离适配效果与基础模型选择的影响。在贪婪解码和约束解码下,通过自动指标和基于LLM的评判评估多选题(MCQA)和开放式问答(OEQA)。对于多选题,CPT+SFT通常得分最高,但相比SFT提升小且常不显著,故SFT是成本效益高的默认选择;对于开放式问答,CPT持续提升重叠度指标,而SFT常降低生成质量,指令微调与CPT+SFT更受基于LLM的评价青睐。跨语言实验表明,法语适配可有效迁移至英文基准。整体提供在计算资源受限下的适配策略选择指南。

原文摘要 · Abstract (English)

The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized domains and languages, yet the effectiveness of domain adaptation strategies remains unclear. We present a study of medical domain adaptation using French medical question-answering (QA) as a case study. We compare continual pretraining (CPT), supervised fine-tuning (SFT), and their combination across three model families, multiple sizes, and three initialization types, explicitly disentangling adaptation effects from base model choice. We evaluate both multiple-choice (MCQA) and open-ended QA (OEQA) under greedy and constrained decoding using automatic metrics and LLM-as-a-Judge evaluation. For MCQA, CPT+SFT most often achieves the best scores, but gains over SFT are small and frequently not statistically significant, making SFT a strong and cost-effective default. For OEQA, CPT consistently improves overlap-based metrics, while SFT often degrades generation quality; instruction tuning and CPT+SFT are preferred by LLM-based evaluation. Cross-lingual experiments further show effective transfer from French adaptation to English benchmarks. Overall, we provide practical guidelines for selecting adaptation strategies under computational constraints.

医疗AI大模型适配法语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。