arXiv:2508.07292cs.AIcs.CL2025-08被引 3

让AI像医生一样一步步验证内镜证据,避免错误累积。

EndoCogniAgent: Closed-Loop Agentic Reasoning with Self-Consistency Validation for Endoscopic Diagnosis

  • 构建闭环智能体,分步采集证据并自我验证一致性。
  • 感知任务准确率85.23%,推理任务临床接受率达71.13%。
  • 适合需要高可靠诊断的医疗AI研究与应用者。

内镜诊断是临床医生逐步获取、比较和验证局部视觉证据以形成结论的迭代过程。当前AI系统未能充分支持此过程,因细粒度证据获取与多步推理耦合不足,导致幻觉证据和错误累积两大失效模式,影响诊断可靠性。本文提出EndoCogniAgent,一种闭环智能体框架,将内镜诊断建模为受控状态更新过程。每轮推理中,中央规划器选择下一步证据获取动作,专用专家工具提取对应观测,自一致性验证机制从知识一致性(与输入图像匹配)和时间一致性(与先前已验证发现一致)两个维度评估观测,再决定是否更新诊断状态。经验证的观测被纳入演化状态以指导后续规划,未充分支持的发现则保留并提供纠正反馈,引导规划器进行额外验证。此外,我们构建了EndoAgentBench,一个面向工作流的基准,包含来自11个内镜数据集的6,132个问答对,用于评估诊断代理在从细粒度视觉感知到高层推理的完整诊断链上的表现。实验表明,EndoCogniAgent在感知任务上平均准确率达85.23%,在推理任务上临床接受率为71.13%;消融分析证实自一致性验证与情景状态维护对性能提升均至关重要。

原文摘要 · Abstract (English)

Endoscopic diagnosis is an iterative process in which clinicians progressively acquire, compare, and verify local visual evidence before reaching a conclusion. Current AI systems do not adequately support this process because fine-grained evidence acquisition and multi-step reasoning remain weakly coupled. This gives rise to two failure modes, hallucinated evidence and uncorrected error accumulation, that undermine diagnostic reliability. We propose EndoCogniAgent, a closed-loop agentic framework that formulates endoscopic diagnosis as a controlled state update process. At each reasoning round, a central planner selects the next evidence acquisition action, specialized expert tools extract the corresponding observation, and a self-consistency validation mechanism examines the observation along two dimensions, knowledge consistency against the input image and temporal consistency with prior validated findings, before updating the diagnostic state. Validated observations are admitted into the evolving state to condition subsequent planning, while insufficiently supported findings are retained with corrective feedback that redirects the planner toward additional verification. We further introduce EndoAgentBench, a workflow-oriented benchmark comprising 6,132 question-answer pairs from 11 endoscopic datasets, designed to evaluate diagnostic agents across a comprehensive diagnostic chain, from fine-grained visual perception to high-level diagnostic reasoning. Experiments show that EndoCogniAgent achieves 85.23\% average accuracy on perception tasks and 71.13\% clinical acceptance rate on reasoning tasks, with ablation analysis confirming that self-consistency validation and episodic state maintenance are individually critical to these gains.

医疗AI闭环推理自一致性内镜诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。