arXiv:2510.21324cs.AIcs.MA2025-10被引 7

用导演式多阶段协作,让AI更可靠地解读胸部X光片。

CXRAgent: Director-Orchestrated Multi-Stage Reasoning for Chest X-Ray Interpretation

  • 设立中央导演协调工具调用与专家团队协作
  • 通过视觉证据验证输出,提升诊断可信度
  • 适合需要高可靠性解释的临床场景使用

胸部X光(CXR)在临床诊断中至关重要,已有多种任务专用和基础模型用于自动解读。然而,这些模型在应对新诊断任务和复杂推理时表现不佳。近期基于大语言模型的智能体成为新范式,通过工具协同、多步推理和团队合作增强能力,但现有智能体通常依赖单一诊断流程,缺乏对工具可靠性的评估机制,限制了适应性与可信度。为此,我们提出CXRAgent,一种由导演统筹的多阶段智能体,包含三个阶段:(1) 工具调用:智能体策略性调用一组CXR分析工具,输出经由基于证据的验证器(EDV)标准化并验证,以视觉证据支撑诊断;(2) 诊断规划:根据任务需求与中间发现制定针对性诊断计划,组建专家团队并明确角色与协作方式,实现自适应协作推理;(3) 协同决策:整合专家团队意见与累积上下文记忆,生成有证据支持的诊断结论。在多个CXR解读任务上的实验表明,CXRAgent表现优异,能提供可视化证据,并有效泛化至不同复杂度的临床任务。代码与数据可在链接获取。

原文摘要 · Abstract (English)

Chest X-ray (CXR) plays a pivotal role in clinical diagnosis, and a variety of task-specific and foundation models have been developed for automatic CXR interpretation. However, these models often struggle to adapt to new diagnostic tasks and complex reasoning scenarios. Recently, LLM-based agent models have emerged as a promising paradigm for CXR analysis, enhancing model's capability through tool coordination, multi-step reasoning, and team collaboration, etc. However, existing agents often rely on a single diagnostic pipeline and lack mechanisms for assessing tools' reliability, limiting their adaptability and credibility. To this end, we propose CXRAgent, a director-orchestrated, multi-stage agent for CXR interpretation, where a central director coordinates the following stages: (1) Tool Invocation: The agent strategically orchestrates a set of CXR-analysis tools, with outputs normalized and verified by the Evidence-driven Validator (EDV), which grounds diagnostic outputs with visual evidence to support reliable downstream diagnosis; (2) Diagnostic Planning: Guided by task requirements and intermediate findings, the agent formulates a targeted diagnostic plan. It then assembles an expert team accordingly, defining member roles and coordinating their interactions to enable adaptive and collaborative reasoning; (3) Collaborative Decision-making: The agent integrates insights from the expert team with accumulated contextual memories, synthesizing them into an evidence-backed diagnostic conclusion. Experiments on various CXR interpretation tasks show that CXRAgent delivers strong performance, providing visual evidence and generalizes well to clinical tasks of different complexity. Code and data are valuable at this \href{https://github.com/laojiahuo2003/CXRAgent/}{link}.

医学影像智能体系统多阶段推理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。