arXiv:2603.20435cs.AI2026-03

通过自我反思迭代修正,提升临床文本结构化提取的逻辑一致性。

Deep reflective reasoning in interdependence constrained structured data extraction from clinical notes for digital health

  • 设计自检自修机制,持续验证变量间逻辑关系与文本一致性。
  • 在三种肿瘤场景中,平均准确率提升10%以上,关键指标最高达94.8%。
  • 适合需要高可靠临床数据的医疗AI研发人员使用。

从临床笔记中提取结构化信息需处理众多相互依赖的变量,其取值存在逻辑约束。现有基于大语言模型(LLM)的抽取流程常因忽略此类依赖而产生临床不一致结果。本文提出深度反思推理框架,通过迭代式自我批判与修正,检查变量间一致性、输入文本及检索到的领域知识,在输出收敛时停止。在三个不同肿瘤应用中评估:(1) 结直肠癌大体描述生成结构化报告(n=217),八项分类变量平均F1从0.828升至0.911,四项数值变量正确率从0.806升至0.895;(2) Ewing肉瘤CD99免疫染色模式识别(n=200),准确率由0.870提升至0.927;(3) 肺癌分期(n=100),分期准确率从0.680升至0.833(pT: 0.842→0.884;pN: 0.885→0.948)。结果表明,深度反思推理可系统性提升在依赖约束下的结构化抽取可靠性,助力生成更一致的机器可操作临床数据集,推动数字健康中的机器学习与数据科学应用。

原文摘要 · Abstract (English)

Extracting structured information from clinical notes requires navigating a dense web of interdependent variables where the value of one attribute logically constrains others. Existing Large Language Model (LLM)-based extraction pipelines often struggle to capture these dependencies, leading to clinically inconsistent outputs. We propose deep reflective reasoning, a large language model agent framework that iteratively self-critiques and revises structured outputs by checking consistency among variables, the input text, and retrieved domain knowledge, stopping when outputs converge. We extensively evaluate the proposed method in three diverse oncology applications: (1) On colorectal cancer synoptic reporting from gross descriptions (n=217), reflective reasoning improved average F1 across eight categorical synoptic variables from 0.828 to 0.911 and increased mean correct rate across four numeric variables from 0.806 to 0.895; (2) On Ewing sarcoma CD99 immunostaining pattern identification (n=200), the accuracy improved from 0.870 to 0.927; (3) On lung cancer tumor staging (n=100), tumor stage accuracy improved from 0.680 to 0.833 (pT: 0.842 -> 0.884; pN: 0.885 -> 0.948). The results demonstrate that deep reflective reasoning can systematically improve the reliability of LLM-based structured data extraction under interdependence constraints, enabling more consistent machine-operable clinical datasets and facilitating knowledge discovery with machine learning and data science towards digital health.

临床信息提取大模型推理医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。