arXiv:2609.04366cs.CL2026-09

用验证循环提升临床文本中结直肠癌早期症状提取的准确性

VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes

论文配图:VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes
图 1 · 摘自论文原文
  • 构建可自我验证的智能流程,逐步修正错误标注
  • 精度从0.764提升至0.849,误报率显著下降
  • 仅1.5%的条目需人工审核,适合医疗场景落地

年轻成人中早期结直肠癌发病率上升,但该人群的警示症状尚无基于证据的随访指南,结构化就诊数据也未能捕捉症状持续时间、背景及家族史等关键信息。本研究开发并评估了一种自动化方法,从自由文本临床记录中提取六类警示症状和家族史风险状态。提出VERGE框架,采用检索增强生成生成初步标签与依据,再通过有限次数的验证-修正循环检查文本依据与临床合理性,直至解决或达到上限,未解决项交由人工复核。在4,033对临床医生标注的文本-发现对上测试,相较于单代理基线,VERGE将精确率从0.764提升至0.849,马修相关系数从0.681升至0.730,在精度与召回间取得均衡提升,98.5%的错误可自主修正,仅1.5%需人工介入。结果表明,受控的验证式工作流可在不牺牲检出能力的前提下减少无效阳性,为年轻患者结直肠癌风险评估提供更可靠、可信的临床语言处理工具。

原文摘要 · Abstract (English)

Early-onset colorectal cancer is increasing among younger adults, yet red-flag symptoms in this age group have no evidence-based guidelines for follow-up testing, and structured encounter data do not capture the detail needed to support early detection and inform follow-up, including symptom duration, context, and fam- ily history, an established colorectal-cancer risk factor. This study aimed to develop and evaluate an automated method for extracting six red-flag symptoms and family-history risk status from free-text clinical notes. We developed VERGE, an agentic workflow in which an initial label and evidence are proposed using retrieval-augmented generation, then passed through a bounded verification- refinement cycle that checks textual grounding and clinical validity, corrects and rechecks a claim until resolved or a limit is reached, and escalates unresolved claims for human review. VERGE was evaluated on 4,033 clinician-labeled note-finding pairs against a single-agent baseline, a rule-based clinical language-processing baseline, and an alternative underlying language model. Compared with the single-agent baseline, VERGE reduced false positive find- ings, improving precision from 0.764 to 0.849 and MCC from 0.681 to 0.730, a balanced gain across the precision-recall trade-off, and resolved most flagged errors autonomously, with human review required for only 1.5 percent of claims. These results indicate that a bounded, verification-based workflow can reduce unnecessary positive findings without sacrificing the ability to detect true ones. This approach offers a path toward more reliable and trustworthy clinical language-processing tools to support colorectal cancer risk assessment in younger patients.

医疗AI文本抽取验证机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。