用智能代理系统从海量癌病记录中自动提取结构化数据,准确率超93%
HARMON-E: Hierarchical Agentic Reasoning for Multimodal Oncology Notes to Extract Structured Data
- 将复杂抽取任务拆解为可自适应的模块化代理任务
- 在40万份病历上实现平均F1达0.93,关键变量超0.95
- 适合临床研究与医疗数据治理场景,大幅降低人工成本
电子健康记录中的非结构化病历包含对癌症治疗决策和研究至关重要的丰富临床信息,但因内容差异大、术语专业、格式不一,可靠提取结构化数据仍具挑战。人工标注虽准确但成本高且难扩展。现有自动化方法多局限于特定场景——或依赖合成数据,或仅处理文档级提取,或孤立关注个别变量(如分期、生物标志物、组织学),难以应对大量病历中存在矛盾信息的患者级综合分析。本研究提出一种基于智能体的框架,将复杂的肿瘤数据抽取任务系统性分解为模块化、可自适应的任务。具体使用大语言模型(LLMs)作为推理代理,结合上下文敏感检索与迭代合成能力,全面、细致地从真实世界肿瘤病历中提取结构化临床变量。在涵盖2,250名癌症患者、超过40万份未结构化临床笔记及扫描PDF报告的大规模数据集上评估,该方法平均F1得分为0.93,103个肿瘤相关变量中有100个超过0.85,关键变量(如生物标志物、药物)均超过0.95。此外,将该智能体系统集成至数据整理流程后,实现0.94的直接人工审核通过率,显著降低标注成本。据我们所知,这是首个在大规模场景下端到端应用基于LLM的智能体进行结构化肿瘤数据抽取的研究。
原文摘要 · Abstract (English)
Unstructured notes within the electronic health record (EHR) contain rich clinical information vital for cancer treatment decision making and research, yet reliably extracting structured oncology data remains challenging due to extensive variability, specialized terminology, and inconsistent document formats. Manual abstraction, although accurate, is prohibitively costly and unscalable. Existing automated approaches typically address narrow scenarios - either using synthetic datasets, restricting focus to document-level extraction, or isolating specific clinical variables (e.g., staging, biomarkers, histology) - and do not adequately handle patient-level synthesis across the large number of clinical documents containing contradictory information. In this study, we propose an agentic framework that systematically decomposes complex oncology data extraction into modular, adaptive tasks. Specifically, we use large language models (LLMs) as reasoning agents, equipped with context-sensitive retrieval and iterative synthesis capabilities, to exhaustively and comprehensively extract structured clinical variables from real-world oncology notes. Evaluated on a large-scale dataset of over 400,000 unstructured clinical notes and scanned PDF reports spanning 2,250 cancer patients, our method achieves an average F1-score of 0.93, with 100 out of 103 oncology-specific clinical variables exceeding 0.85, and critical variables (e.g., biomarkers and medications) surpassing 0.95. Moreover, integration of the agentic system into a data curation workflow resulted in 0.94 direct manual approval rate, significantly reducing annotation costs. To our knowledge, this constitutes the first exhaustive, end-to-end application of LLM-based agents for structured oncology data extraction at scale
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。