AI通过迭代查证证据,精准推荐药物治疗方案。
An AI agent for treatment reasoning over a biomedical tool universe

- 构建212个生物医学工具库,用强化学习训练AI逐步查证信息。
- 在3168个用药推理任务中准确率达94.7%,比GPT-5高17.8点。
- 适用于罕见病、重症心血管与感染疾病等复杂临床场景。
治疗决策依赖于对疾病背景、共病、用药禁忌及不断更新的生物医学知识的综合判断,具有高度迭代性:候选方案需反复权衡约束条件,随新证据动态调整,并基于可验证来源。本文提出ATHENA-R1,一个覆盖自1939年以来所有FDA批准药物的治疗推理AI代理,基于212个生物医学工具构建的宇宙进行强化学习训练。该模型每一步识别缺失信息,选择并运行相关工具,融合新证据。为避免依赖人工标注,采用两级自学习框架:多智能体系统生成工具、任务和推理轨迹用于监督微调;再通过科学反馈奖励机制(证据获取、工具使用合理性、逻辑无冗余)进行强化学习。在涵盖3,168个药物推理任务和456个患者治疗案例的五个基准测试中,ATHENA-R1在开放式药物推理任务中达到94.7%准确率,在治疗推理任务中达82.9%,分别领先GPT-5 17.8和10.7个百分点。28家罕见病组织专家盲评显示,其在所有指标上均优于参考模型;医生在复杂住院心血管与感染性疾病案例中也给予积极评价。其生成的不良事件假设在包含540万患者电子健康记录的数据中验证,校正后优势比达1.48–1.84,而阴性对照组未见升高。由于治疗推理要求在结论前明确所需证据,传统上难以实现,我们证明其可被重构为可学习的迭代证据获取过程,强化学习可有效训练AI完成。
原文摘要 · Abstract (English)
Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an appropriate therapy. It is inherently iterative: candidates are weighed against many constraints, revised as evidence emerges, and grounded in verifiable sources. Here we introduce ATHENA-R1, an AI agent for treatment reasoning across all FDA approved drugs since 1939, trained by reinforcement learning over a universe of 212 biomedical tools. At each step it identifies missing information, selects and runs relevant tools, and incorporates the evidence. To train it without human-annotated traces, we build a two-level self-learning framework: multi-agent systems construct the tools, tasks, and reasoning trajectories for supervised fine-tuning, then reinforcement learning with scientific feedback rewards reasoning quality (evidence gathering, grounded tool use, logical non-redundancy). Across five benchmarks of 3,168 drug reasoning tasks and 456 patient treatment cases, ATHENA-R1 outperforms language models and tool-use systems, reaching 94.7% accuracy on open-ended drug reasoning and 82.9% on treatment reasoning, 17.8 and 10.7 points above GPT-5. In blinded evaluations by experts from 28 rare disease organizations, it is preferred over reference models on all criteria, and physicians rated it favorably on complex hospitalized cardiovascular and infectious-disease cases. Adverse-event hypotheses it generated, tested in electronic health records from 5.4 million patients, reached adjusted odds ratios of 1.48-1.84, with no elevation among negative controls. Because it requires knowing what evidence to seek before concluding, treatment reasoning has long been hard for AI; we show it can be reframed as a learnable process of iterative evidence gathering that reinforcement learning can train AI to perform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。