arXiv:2607.12051cs.CL2026-07

用智能体系统生成乳腺癌治疗建议,效果尚可但仍有明显错误。

Agentic systems for breast cancer treatment recommendations

论文配图:Agentic systems for breast cancer treatment recommendations
图 1 · 摘自论文原文
  • 构建多智能体系统,结合工具调用与自主子代理,评估治疗建议生成能力。
  • 最佳模型在72个真实病例中得分0.594,但不同阶段表现差异大。
  • 适合临床辅助决策研究者,不适用于无人监督的医疗场景。

大型语言模型(LLMs)在临床决策支持中日益受到关注,但在复杂肿瘤治疗规划中的可靠性仍不明确。本研究基于72例分期为I至IV期的真实乳腺癌病例,采用1,147个通过非对称信息评分生成法(AIRG)创建的案例专用评分标准,评估了智能体式LLM系统在治疗建议生成中的表现。共比较七种方案,包括单模型基线、工具增强系统及具备事实核查与自主子代理生成能力的多智能体架构。最佳配置为Claude Opus 4.8配合D&C+SA流水线,全局得分为0.594 ± 0.025。工具使用与提升智能体自主性效果不一,在某些场景提升性能,其他场景则导致下降。性能随临床领域和疾病分期变化,且由肿瘤科医生主导的错误分析揭示了持续存在的临床相关缺陷,包括错误或遗漏建议、推理瑕疵、引用错误、过时陈述及过度自信。结果表明,智能体式LLM系统虽能生成具有临床意义的乳腺癌治疗建议,但仍不足以用于无监督临床应用。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning remains unclear. We evaluated agentic LLM systems for breast cancer treatment recommendation generation using 72 real clinical cases across stages I to IV and 1,147 case-specific rubrics generated through Asymmetric Information Rubric Generation (AIRG), in which the rubric generator had access to real clinical decisions unavailable to the evaluated models. Seven pipelines were compared, including single-LLM baselines, tool-augmented systems, and multi-agent architectures with fact checking and autonomous subagent spawning. The best-performing configuration, Claude Opus 4.8 with the D&C+SA pipeline, achieved a global score of 0.594 $\pm$ 0.025. Tool use and increased agent autonomy had mixed effects, improving performance in some settings but degrading it in others. Performance varied by clinical domain and disease stage, and oncologist-led error analysis revealed persistent clinically relevant failures, including incorrect or missing recommendations, flawed justifications, citation errors, outdated claims, and overconfidence. These findings suggest that agentic LLM systems can generate clinically relevant breast cancer recommendations, but remain insufficient for unsupervised clinical use.

智能体乳腺癌医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。