用提示工程提升大模型辅助药物流行病学研究设计,通用模型表现优于专业医学模型。
Employing General-Purpose and Biomedical Large Language Models with Advanced Prompt Engineering for Pharmacoepidemiologic Study Design
- 采用分步提示和主动提示策略,提升大模型推理能力。
- GPT-4o在8个问题中达最高相关性得分(中位数4分),优于医学专用模型。
- 医学大模型常缺乏充分论证,且对术语编码映射能力弱。
大型语言模型(LLMs)在自动化支持药物流行病学研究设计方面具有潜力,但其可靠性尚未充分验证。本研究对比了通用大模型(GPT-4o 和 DeepSeek-R1)与生物医学微调模型(QuantFactory/Bio-Medical-Llama-3-8B-GGUF 与 Irathernotsay/qwen2-1.5B-medical_qa-Finetune),使用2018–2024年来自HMA-EMA目录与哨兵系统共46份研究协议进行评估。通过最少到最多(LTM)与主动提示策略,考察模型在相关性、论证逻辑及多编码系统术语匹配方面的表现。结果显示,搭配LTM提示的GPT-4o与DeepSeek-R1在相关性和论证逻辑上得分最高,其中GPT-4o-LTM在9个问题中的8个达到中位相关性4分;而医学专用模型整体相关性较低,论证常不充分。所有模型在术语编码映射方面均表现有限,但LTM策略显著提升推理稳定性。结论:当前现成通用大模型在药物流行病学设计支持上优于生物医学专用模型,提示策略对性能影响显著。
原文摘要 · Abstract (English)
Background: The potential of large language models (LLMs) to automate and support pharmacoepidemiologic study design is an emerging area of interest, yet their reliability remains insufficiently characterized. General-purpose LLMs often display inaccuracies, while the comparative performance of specialized biomedical LLMs in this domain remains unknown. Methods: This study evaluated general-purpose LLMs (GPT-4o and DeepSeek-R1) versus biomedically fine-tuned LLMs (QuantFactory/Bio-Medical-Llama-3-8B-GGUF and Irathernotsay/qwen2-1.5B-medical_qa-Finetune) using 46 protocols (2018-2024) from the HMA-EMA Catalogue and Sentinel System. Performance was assessed across relevance, logic of justification, and ontology-code agreement across multiple coding systems using Least-to-Most (LTM) and Active Prompting strategies. Results: GPT-4o and DeepSeek-R1 paired with LTM prompting achieved the highest relevance and logic of justification scores, with GPT-4o-LTM reaching a median relevance score of 4 in 8 of 9 questions for HMA-EMA protocols. Biomedical LLMs showed lower relevance overall and frequently generated insufficient justification. All LLMs demonstrated limited proficiency in ontology-code mapping, although LTM provided the most consistent improvements in reasoning stability. Conclusion: Off-the-shelf general-purpose LLMs currently offer superior support for pharmacoepidemiologic design compared to biomedical LLMs. Prompt strategy strongly influenced LLM performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。