用大模型分析病历,为肝癌患者精准分层并推荐治疗方案。
Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance

- 基于临床推理的LLM读取电子病历,输出风险评分与治疗建议。
- 在6668例患者中,按建议治疗可延长至51个月生存期。
- 医生评价其推理可信,能提升诊疗效率与准确性。
肝细胞癌(HCC)是常见恶性肿瘤和癌症死亡主因。现有指南与分期系统分类粗糙,常忽略同一分期内异质性及电子病历(EMRs)中的临床背景。我们提出HCC-STAR(肝细胞癌分型、治疗与预后),一个与临床对齐的大语言模型,可读取常规EMR文本,联合输出基于风险评分的分期、符合指南的治疗排序及其证据支持理由,以及个体化生存预测。我们从SEER数据库收集约3万例HCC病例,并通过医师验证的提示增强工作流,生成类病历叙事训练数据。在此语料上,我们构建了知识对齐的推理框架,采用可逐步验证的复合奖励优化,超越文本层面的记忆。在来自中国12家医院的多中心队列(6,668名患者)中,HCC-STAR在治疗推荐与风险分层方面表现优于临床指南及主流模型(包括GPT-5与Gemini-2.5 Pro)。假设总生存分析显示,遵循HCC-STAR建议的患者中位生存期为51个月,显著高于BCLC(29个月)与CNLC(32个月)。在以医生为中心的评估中,盲法肝胆专科医师认为其推理与证据依据可信。该模型在治疗准确率上超过住院医师与主治医师,并作为助手帮助医生更快做出更准确决策。结果表明,HCC-STAR是肝癌风险分层与精准治疗中可靠且可验证的决策支持系统。
原文摘要 · Abstract (English)
Hepatocellular carcinoma (HCC) is a common malignancy and a leading cause of cancer-related mortality. Current guidelines and staging systems provide coarse categories, but often miss within-stage heterogeneity and the clinical context in electronic medical records (EMRs). We present HCC-STAR (Hepatocellular Carcinoma Staging, Treatment And pRognosis), a clinically aligned large language model that reads routine EMR narratives and jointly outputs risk score-based staging, ranked guideline-consistent treatments with evidence-based rationales, and individualized survival estimates. We curated about 30,000 HCC cases from SEER and expanded them into EMR-style narrative training data using a clinician-validated, prompt-based augmentation workflow. On this corpus, we developed a knowledge-aligned reasoning framework optimized with a step-verifiable composite reward, moving beyond text-level memorization of clinical guidelines. In a multi-center cohort of 6,668 patients from 12 hospitals in China, HCC-STAR achieved state-of-the-art performance in treatment recommendation and risk stratification compared with clinical guidelines and competitive models, including GPT-5 and Gemini-2.5 Pro. Hypothetical overall-survival analysis showed a median survival of 51 months under adherence to HCC-STAR recommendations, compared with 29 and 32 months under BCLC and CNLC. In clinician-centric evaluations, blinded hepatobiliary specialists rated HCC-STAR's reasoning and evidence-based justifications as trustworthy. The model surpassed resident and attending physicians in treatment accuracy and helped physicians make more accurate decisions faster when used as an assistant. These findings support HCC-STAR as a reliable and verifiable decision-support system for risk stratification and precision therapy in HCC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。