让AI像医生一样反复推敲治疗方案,更精准安全。
TheraAgent: Self-Improving Therapeutic Agent for Precise and Comprehensive Treatment Planning

- 用生成-判断-优化循环替代一次性输出,模仿医生思考过程。
- 在HealthBench上准确率和完整性领先,专家评估胜率达86%。
- 专设临床评价模块,确保方案符合医疗规范,适合临床辅助场景。
制定治疗方案本质上是复杂的推理与迭代优化任务,而非简单生成问题。现有大语言模型主要依赖一次性输出,缺乏显式验证,可能导致方案粗糙、不完整甚至存在安全隐患。为此,我们提出TheraAgent,一种以迭代生成-判断-优化流程为核心的智能体框架,模拟人类专家反复修正治疗方案的真实思维过程,将粗略不完整的初稿逐步优化为精确、全面且更安全的治疗方案。为强化判断环节,我们引入TheraJudge——一个嵌入推理流程中的治疗专用评估模块,用于强制执行临床标准。实验表明,TheraAgent在HealthBench上达到当前最优性能,兼具高准确率与完整性;专家评估中,其胜率高达86%,在靶向性与风险控制方面表现优异。此外,TheraJudge与HealthBench评估结果高度一致,验证了框架的可靠性。
原文摘要 · Abstract (English)
Formulating a treatment plan is inherently a complex reasoning and refinement task rather than a simple generation problem. However, existing large language models (LLMs) mainly rely on one-shot output without explicit verification, which may result in rough, incomplete, and potentially unsafe treatment plans. To address these limitations, we propose TheraAgent, an agentic framework that replaces one-shot generation with an iterative generate-judge-refine pipeline. By mirroring the actual reasoning process of human experts who iteratively revise treatment plans, our framework progressively transforms coarse and incomplete drafts into precise, comprehensive, and safer therapeutic regimens. To facilitate the critical judge component, we introduce TheraJudge, a treatment-specific evaluation module integrated into the inference loop to enforce clinical standards. Experiments show TheraAgent achieves state-of-the-art results on HealthBench, leading in Accuracy and Completeness. In expert evaluations, it attains an 86% win rate against physicians, with superior Targeting and Harm Control. Moreover, the highly agreement between TheraJudge and HealthBench evaluations confirms the reliability of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。