优化提示与轻量微调,让小模型在临床推理上媲美大模型。
Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies
- 设计四种提示层级,测试其对临床推理的影响
- LoRA微调使性能提升8-12点,接近GPT-4o-mini
- 提示结构影响最大,占性能差异44%,需针对性评估
大型语言模型(LLM)的提示策略与微调技术对其推理能力有显著影响,但在临床自然语言推理(NLI)领域的效果仍不明确。本研究首次在受控环境下评估提示结构与高效微调联合对临床NLI性能的作用。我们测试了四类提示策略以激发不同抽象层级的推理,并利用前沿模型通过低秩适应(LoRA)将多步推理能力蒸馏至40亿参数的小模型中。在NLI4CT基准上,不同提示类型单独可解释高达44%的宏平均F1方差;LoRA微调带来8至12的F1提升,输出一致性超97%,且与GPT-4o-mini的差距缩小至7.1%以内。在医学文本推理与临床试验数据集上的泛化实验显示,LoRA在75%的模型中表现更优。结果表明:(i) 提示结构是临床推理性能的主要驱动因素;(ii) 搭配优质提示与LoRA的小模型可媲美前沿系统;(iii) 需基于推理类型进行评估,才能发现提示带来的权衡。研究为高效、可信的临床NLP系统提供了新思路。
原文摘要 · Abstract (English)
Recent works on large language models (LLMs) have demonstrated the impact of prompting strategies and fine-tuning techniques on their reasoning capabilities. Yet, their effectiveness on clinical natural language inference (NLI) remains underexplored. This study presents the first controlled evaluation of how prompt structure and efficient fine-tuning jointly shape model performance in clinical NLI. We inspect four classes of prompting strategies to elicit reasoning in LLMs at different levels of abstraction, and evaluate their impact on a range of clinically motivated reasoning types. For each prompting strategy, we construct high-quality demonstrations using a frontier model to distil multi-step reasoning capabilities into smaller models (4B parameters) via Low-Rank Adaptation (LoRA). Across different language models fine-tuned on the NLI4CT benchmark, we found that prompt type alone accounts for up to 44% of the variance in macro-F1. Moreover, LoRA fine-tuning yields consistent gains of +8 to 12 F1, raises output alignment above 97%, and narrows the performance gap to GPT-4o-mini to within 7.1%. Additional experiments on reasoning generalisation reveal that LoRA improves performance in 75% of the models on MedNLI and TREC Clinical Trials Track. Overall, these findings demonstrate that (i) prompt structure is a primary driver of clinical reasoning performance, (ii) compact models equipped with strong prompts and LoRA can rival frontier-scale systems, and (iii) reasoning-type-aware evaluation is essential to uncover prompt-induced trade-offs. Our results highlight the promise of combining prompt design and lightweight adaptation for more efficient and trustworthy clinical NLP systems, providing insights on the strengths and limitations of widely adopted prompting and parameter-efficient techniques in highly specialised domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。