用真实医学工具增强大模型,让AI更准地预测疾病风险。
RiskAgent: Synergizing Language Models with Validated Tools for Evidence-Based Risk Prediction
- 将大模型与数百个循证医学工具结合,避免幻觉。
- 在多种疾病风险预测中表现优于现有方法。
- 适合医疗决策、临床研究等需要可靠推理的场景。
大型语言模型(LLMs)在医学考试中已达到与人类专家相当的水平,但在复杂临床决策中仍面临挑战,因其需深入医学知识理解,而当前多数研究局限于标准化试题场景。主流方法为微调模型,但需大量数据和算力,且易产生‘幻觉’。本文提出RiskAgent,通过整合数百个经验证的循证医学临床决策工具,实现通用且可信的风险预测。实验表明,RiskAgent在跨疾病、多场景的临床风险预测任务中表现优异,且在外部MedCalc-Bench数据集上展现出强大的工具学习泛化能力,在MedQA、MedMCQA和MMLU三个基准上也具备出色的医学推理与问答能力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) achieve competitive results compared to human experts in medical examinations. However, it remains a challenge to apply LLMs to complex clinical decision-making, which requires a deep understanding of medical knowledge and differs from the standardized, exam-style scenarios commonly used in current efforts. A common approach is to fine-tune LLMs for target tasks, which, however, not only requires substantial data and computational resources but also remains prone to generating `hallucinations'. In this work, we present RiskAgent, which synergizes language models with hundreds of validated clinical decision tools supported by evidence-based medicine, to provide generalizable and faithful recommendations. Our experiments show that RiskAgent not only achieves superior performance on a broad range of clinical risk predictions across diverse scenarios and diseases, but also demonstrates robust generalization in tool learning on the external MedCalc-Bench dataset, as well as in medical reasoning and question answering on three representative benchmarks, MedQA, MedMCQA, and MMLU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。