arXiv:2607.06435cs.AI2026-07被引 1

通过可配置多代理系统实现病理报告中临床信息的精准提取与证据溯源。

Trust but Verify:Evidence-Linked Multi-Agent Clinical Information Extraction in Pathology

论文配图:Trust but Verify:Evidence-Linked Multi-Agent Clinical Information Extraction in Pathology
图 1 · 摘自论文原文
  • 分离字段定义与模型,支持可配置工作流
  • 98.61%准确率,所有正确预测均有原文证据支持
  • 适合需要可解释性与医生审核的临床信息抽取场景

从病理报告中提取临床特征极具挑战,因相关信息可能分散于编码和文本字段,且依赖样本归属、否定判断、辅助发现及诊断背景。本研究回顾性评估了可配置的NimbleMind多代理系统(nMAS),该系统将临床医生定义的字段规范与抽取模型解耦,并返回带来源证据的报告级预测。研究使用来自新加坡的54份模拟胃活检报告,涵盖四个二值目标字段,共216个特征-病例决策。nMAS正确分类213/216个决策(98.61%),所有正确预测的证据跨度均在原文中完全一致。三个错误均出现在两个依赖上下文的幽门螺杆菌相关字段,需处理否定或诊断归属。单一模型的UMA式对比方法表现相近,错误模式一致。结果未显示多代理架构在预测性能上的优势。nMAS的核心贡献在于工作流整合、可配置字段规范、基于复杂度的路由、报告级聚合以及医生可审查流程中的源文本验证。未来需开展更大规模跨机构研究,评估泛化能力、语义证据质量、适配成本与医生验证时间。

原文摘要 · Abstract (English)

Clinical feature extraction from pathology reports is challenging because relevant evidence may be distributed across coded and narrative fields and depend on specimen attribution, negation, ancillary findings, and diagnostic context. We retrospectively evaluated the NimbleMind Multi-Agent System (nMAS), a configurable workflow that separates clinician-defined field specifications from extraction models and returns report-level predictions with source-linked evidence. The study included 54 dummy gastric biopsy pathology reports from Singapore and four binary target fields, yielding 216 feature-case decisions. nMAS correctly classified 213 of 216 decisions (98.61\%), and all evidence spans associated with correct predictions occurred verbatim in the corresponding source reports. All three errors occurred in the two context-dependent \textit{H. pylori}-related fields requiring negation handling or diagnostic attribution. A single-model UMA-style comparator produced the similar label-level performance and error pattern. These findings do not demonstrate predictive superiority for the multi-agent architecture.Rather, the contribution of nMAS lies in workflow integration and traceability through configurable field specifications, complexity-based routing, report-level aggregation, and source-text validation within a clinician-reviewable workflow. Larger multi-institutional studies should assess generalizability, semantic evidence quality, adaptation effort, and clinician verification time.

信息抽取多代理系统病理分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。