用优化检索提升开源医学大模型性能,成本更低效果更好
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
- 通过优化上下文检索增强开源医学大模型
- 在MedQA上达到顶尖准确率,成本仅为商用模型的几分之一
- 开源工具包助力医疗AI开发,适合研究与临床应用
本研究通过优化上下文检索,提升开源大型语言模型在医疗领域的性价比与性能。实验表明,该方法在医学问答任务中实现当前最优准确率,且成本仅为商用模型的极小部分,在MedQA基准上显著改善了成本-准确率帕累托前沿。主要贡献包括:(1) 提出OpenMedQA新基准,揭示开放问答与选择题格式间的性能差距;(2) 构建可复现的上下文检索优化流程;(3) 开源提示工程工具与思维链/思维树数据库(Prompt Engine, CoT/ToT/Thinking databases),支持医疗AI研发。通过改进检索与评估方式,推动更经济可靠的医疗大模型应用。
原文摘要 · Abstract (English)
This study leverages optimized context retrieval to enhance open-source Large Language Models (LLMs) for cost-effective, high performance healthcare AI. We demonstrate that this approach achieves state-of-the-art accuracy on medical question answering at a fraction of the cost of proprietary models, significantly improving the cost-accuracy Pareto frontier on the MedQA benchmark. Key contributions include: (1) OpenMedQA, a novel benchmark revealing a performance gap in open-ended medical QA compared to multiple-choice formats; (2) a practical, reproducible pipeline for context retrieval optimization; and (3) open-source resources (Prompt Engine, CoT/ToT/Thinking databases) to empower healthcare AI development. By advancing retrieval techniques and QA evaluation, we enable more affordable and reliable LLM solutions for healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。