arXiv:2509.07325cs.LG2025-09被引 1

用大模型自动生成肺癌治疗指南推荐,降低医生负担。

CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement Estimation

  • 构建121例肺癌患者纵向数据集,由专家标注治疗路径。
  • 模型生成建议与专家标注相关性高(斯皮尔曼r=0.88)。
  • 融合人类标注与模型一致性,提升推荐可信度和可解释性。

美国国立综合癌症网络(NCCN)提供基于证据的癌症治疗指南。将复杂的患者情况转化为符合指南的治疗建议耗时长、需专业经验,且易出错。大语言模型(LLM)的发展有望缩短生成时间并提高准确性。本文提出一种基于LLM代理的方法,为非小细胞肺癌(NSCLC)患者自动生成符合指南的治疗轨迹。贡献有三:其一,构建包含121例患者临床记录、诊断结果与病史的纵向数据集,由资深肿瘤科医生逐项标注对应NCCN指南路径;其二,证明现有LLM具备领域知识,可生成高质量代理基准,与专家标注高度一致(斯皮尔曼相关系数r=0.88,RMSE=0.08);其三,提出混合方法,结合昂贵的人工标注与模型一致性信息,建立预测指南的代理框架及验证准确性的元分类器,输出校准置信度分数(AUROC=0.800),支持性能权衡与监管合规。该研究建立了一套兼顾准确性、可解释性与合规性的临床可行指南遵循系统,降低标注成本,推动自动化临床决策支持落地。

原文摘要 · Abstract (English)

The National Comprehensive Cancer Network (NCCN) provides evidence-based guidelines for cancer treatment. Translating complex patient presentations into guideline-compliant treatment recommendations is time-intensive, requires specialized expertise, and is prone to error. Advances in large language model (LLM) capabilities promise to reduce the time required to generate treatment recommendations and improve accuracy. We present an LLM agent-based approach to automatically generate guideline-concordant treatment trajectories for patients with non-small cell lung cancer (NSCLC). Our contributions are threefold. First, we construct a novel longitudinal dataset of 121 cases of NSCLC patients that includes clinical encounters, diagnostic results, and medical histories, each expertly annotated with the corresponding NCCN guideline trajectories by board-certified oncologists. Second, we demonstrate that existing LLMs possess domain-specific knowledge that enables high-quality proxy benchmark generation for both model development and evaluation, achieving strong correlation (Spearman coefficient r=0.88, RMSE = 0.08) with expert-annotated benchmarks. Third, we develop a hybrid approach combining expensive human annotations with model consistency information to create both the agent framework that predicts the relevant guidelines for a patient, as well as a meta-classifier that verifies prediction accuracy with calibrated confidence scores for treatment recommendations (AUROC=0.800), a critical capability for communicating the accuracy of outputs, custom-tailoring tradeoffs in performance, and supporting regulatory compliance. This work establishes a framework for clinically viable LLM-based guideline adherence systems that balance accuracy, interpretability, and regulatory requirements while reducing annotation costs, providing a scalable pathway toward automated clinical decision support.

癌症诊疗大模型应用临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。