让AI像医生一样看胸片,边看边推理,还能自我检查纠错。
RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows
- 用多个专业AI agent协作,模仿医生读片流程
- 融合图像与文字信息,输出有视觉依据的诊断理由
- 能发现不同工具间的矛盾并自动验证,适合临床可信场景
智能体系统通过多智能体协作和工具使用,有望解决复杂临床任务。然而,现有胸片(CXR)解读方法仍存在三大不足:(i) 推理过程缺乏临床可解释性且不符合指南,仅是工具输出的简单拼接;(ii) 多模态证据融合不足,导致解释仅为文本描述,缺乏视觉支撑;(iii) 很少检测或解决跨工具不一致问题,缺乏系统性验证机制。为此,我们提出RadAgents,一种结合临床先验知识与任务感知多模态推理的多智能体框架,将放射科医生的工作流程编码为模块化、可审计的处理管线。同时,集成视觉定位与多模态检索增强技术,实现上下文冲突的验证与修复,最终输出更可靠、透明且符合临床实践的诊断结果。
原文摘要 · Abstract (English)
Agentic systems offer a potential path to solve complex clinical tasks through collaboration among specialized agents, augmented by tool use and external knowledge bases. Nevertheless, for chest X-ray (CXR) interpretation, prevailing methods remain limited: (i) reasoning is frequently neither clinically interpretable nor aligned with guidelines, reflecting mere aggregation of tool outputs; (ii) multimodal evidence is insufficiently fused, yielding text-only rationales that are not visually grounded; and (iii) systems rarely detect or resolve cross-tool inconsistencies and provide no principled verification mechanisms. To bridge the above gaps, we present RadAgents, a multi-agent framework that couples clinical priors with task-aware multimodal reasoning and encodes a radiologist-style workflow into a modular, auditable pipeline. In addition, we integrate grounding and multimodal retrieval-augmentation to verify and resolve context conflicts, resulting in outputs that are more reliable, transparent, and consistent with clinical practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。