GatorOnco用智能体技术生成结直肠癌治疗方案,表现媲美专家。
An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer
- 构建智能体架构,动态融合最新临床指南进行推理。
- 在盲评中显著优于开源模型,多项评分达专家水平。
- 生成方案更清晰完整,适合临床辅助决策场景。
精准肿瘤学的治疗规划需整合多样患者信息与快速更新的临床指南以确保符合指南的照护。尽管大语言模型在诊断任务中展现潜力,但其在高风险治疗规划中的应用受限于复杂推理、遵循时效性指南及安全性问题。本研究提出GatorOnco,一种用于结直肠癌(CRC)治疗规划的智能体式生成大模型。GatorOnco基于总计2820亿个标记的生物医学文本训练,其中包括来自UF Health的1660亿个标记的医疗系统级临床文本。采用领域适应方法,结合预训练、模型合并、两阶段后训练及基于智能体的强化学习。通过智能体检索增强生成(RAG)机制,动态将时效性临床指南融入推理过程。由五位UF Health肿瘤科医生进行的盲法随机临床评估显示,GatorOnco显著优于开源LLMs(P < 0.01),性能接近专家水平。相比专家,GatorOnco在可读性(4.46 vs. 4.19,P < 0.01)和完整性(3.91 vs. 3.52,P < 0.01)上得分更高,正确性(4.09 vs. 4.11,P = 0.921)、时效性(4.04 vs. 3.98,P = 0.478)和安全性(4.22 vs. 4.22,P = 0.999)无统计差异。结果表明,将智能体推理与大规模领域适配相结合,有助于弥合生成式AI在高风险癌症治疗规划中的差距。
原文摘要 · Abstract (English)
Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure guideline-concordant care. While large language models (LLMs) show promise in many diagnostic tasks, their adoption for high-stakes treatment planning is hindered by complex reasoning, adherence to timely clinical guidelines, and safety concerns. In this study, we present GatorOnco, an agentic LLM for colorectal cancer (CRC) treatment planning. GatorOnco is developed using a total of 282 billion tokens of biomedical text, including healthcare system-scale clinical text comprising 166 billion tokens from UF Health. We implemented a domain-adaptation method that integrates pre-training, model merging, a two-stage post-training approach, and agent-based reinforcement learning. An agentic retrieval-augmented generation (RAG) approach dynamically integrates time-sensitive clinical guidelines into the reasoning process. In a blind, randomized clinical evaluation conducted by five UF Health oncologists, GatorOnco significantly outperformed open-source LLMs (P < 0.01) and achieved expert-level performance comparable to UF Health oncologists. Compared with expert oncologists, GatorOnco received significantly higher ratings for readability (4.46 vs. 4.19, P < 0.01) and completeness (3.91 vs. 3.52, P < 0.01), while showing statistically comparable performance in correctness (4.09 vs. 4.11, P = 0.921), currency (4.04 vs. 3.98, P = 0.478), and safety (4.22 vs. 4.22, P = 0.999). These findings demonstrate that integrating agentic reasoning with large-scale domain adaptation can help bridge the gap for generative AI in high-stakes cancer treatment planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。