arXiv:2509.06602cs.LGcs.AI2025-09被引 9

用AI自动生成肿瘤会诊患者摘要,准确率达94%

Demo: Healthcare Agent Orchestrator (HAO) for Patient Summarization in Molecular Tumor Boards

  • 构建多智能体系统,由LLM协调完成临床信息整合
  • 捕捉94%高重要性信息,评估召回率达0.84
  • 无需共享数据即可本地化评估,适合医院部署

分子肿瘤会诊(MTBs)是多学科专家协作评估复杂病例以制定最优治疗方案的场合。患者摘要作为核心环节,通常由肿瘤科医生或助理手动整理病历信息,但该过程耗时、主观且易遗漏关键内容。为此,我们提出医疗智能体编排器(HAO),一种基于大语言模型的AI系统,通过多智能体工作流自动生成准确全面的患者摘要。由于摘要在风格、顺序、同义词和表述上存在差异,传统评估方法难以衡量简洁性与完整性。为此,我们设计了TBFact——一种“模型即裁判”框架,用于评估生成摘要的完整性和简洁性。基于去标识化的肿瘤会诊数据集,我们使用TBFact评估了患者病史智能体,结果显示其捕获了94%的高重要性信息(含部分蕴含),在严格蕴含标准下达到0.84的召回率。此外,TBFact支持无需数据共享的本地化评估,使机构可安全部署。HAO与TBFact共同为MTBs提供可靠、可扩展的智能支持。

原文摘要 · Abstract (English)

Molecular Tumor Boards (MTBs) are multidisciplinary forums where oncology specialists collaboratively assess complex patient cases to determine optimal treatment strategies. A central element of this process is the patient summary, typically compiled by a medical oncologist, radiation oncologist, or surgeon, or their trained medical assistant, who distills heterogeneous medical records into a concise narrative to facilitate discussion. This manual approach is often labor-intensive, subjective, and prone to omissions of critical information. To address these limitations, we introduce the Healthcare Agent Orchestrator (HAO), a Large Language Model (LLM)-driven AI agent that coordinates a multi-agent clinical workflow to generate accurate and comprehensive patient summaries for MTBs. Evaluating predicted patient summaries against ground truth presents additional challenges due to stylistic variation, ordering, synonym usage, and phrasing differences, which complicate the measurement of both succinctness and completeness. To overcome these evaluation hurdles, we propose TBFact, a ``model-as-a-judge'' framework designed to assess the comprehensiveness and succinctness of generated summaries. Using a benchmark dataset derived from de-identified tumor board discussions, we applied TBFact to evaluate our Patient History agent. Results show that the agent captured 94% of high-importance information (including partial entailments) and achieved a TBFact recall of 0.84 under strict entailment criteria. We further demonstrate that TBFact enables a data-free evaluation framework that institutions can deploy locally without sharing sensitive clinical data. Together, HAO and TBFact establish a robust foundation for delivering reliable and scalable support to MTBs.

医疗AI多智能体摘要生成肿瘤会诊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。