EchoAgent让AI像心脏超声医生一样看、测、想,实现全流程自动分析。
EchoAgent: Towards Reliable Echocardiography Interpretation with "Eyes","Hands" and "Minds"

- 构建专属心脏超声知识库,让AI具备专业推理能力。
- 能自动识别48种心腔视图,完成分割与定量测量,准确率达80%。
- 适合临床辅助诊断,尤其适用于缺乏经验的医生或远程医疗。
可靠的心脏超声(Echo)解读对评估心功能至关重要,需要同步运用视觉观察(眼睛)、手动测量(手)和专家知识学习与推理(大脑)。现有方法多聚焦于单一技能组合,如‘眼-手’或‘眼-脑’,难以保障临床可靠性。为此,我们提出EchoAgent,一个面向端到端心脏超声解读的智能体系统,实现‘眼-手-脑’协同工作,学习、观察、操作与推理如同心脏超声科医生。首先,设计基于专业知识的认知引擎,将权威超声指南结构化为知识库,构建定制化‘大脑’。其次,开发分层协作工具包,赋予系统‘眼-手’能力,可自动解析视频流、识别心腔视图、执行解剖分割与定量测量。第三,将感知的多模态证据与专属知识库融合,通过协同推理枢纽生成可解释推断。在CAMUS和MIMIC-EchoQA数据集上评估,覆盖14个心脏解剖区域的48种不同视图。实验表明,EchoAgent在多种结构分析中表现最优,整体准确率高达80.00%。该系统首次实现单一模型集成学习、观察、操作与推理能力,极大提升心脏超声解读的可靠性。
原文摘要 · Abstract (English)
Reliable interpretation of echocardiography (Echo) is crucial for assessing cardiac function, which demands clinicians to synchronously orchestrate multiple capabilities, including visual observation (eyes), manual measurement (hands), and expert knowledge learning and reasoning (minds). While current task-specific deep-learning approaches and multimodal large language models have demonstrated promise in assisting Echo analysis through automated segmentation or reasoning, they remain focused on restricted skills, i.e., eyes-hands or eyes-minds, thereby limiting clinical reliability and utility. To address these issues, we propose EchoAgent, an agentic system tailored for end-to-end Echo interpretation, which achieves a fully coordinated eyes-hands-minds workflow that learns, observes, operates, and reasons like a cardiac sonographer. First, we introduce an expertise-driven cognition engine where our agent can automatically assimilate credible Echo guidelines into a structured knowledge base, thus constructing an Echo-customized mind. Second, we devise a hierarchical collaboration toolkit to endow EchoAgent with eyes-hands, which can automatically parse Echo video streams, identify cardiac views, perform anatomical segmentation, and quantitative measurement. Third, we integrate the perceived multimodal evidence with the exclusive knowledge base into an orchestrated reasoning hub to conduct explainable inferences. We evaluate EchoAgent on CAMUS and MIMIC-EchoQA datasets, which cover 48 distinct echocardiographic views spanning 14 cardiac anatomical regions. Experimental results show that EchoAgent achieves optimal performance across diverse structure analyses, yielding overall accuracy of up to 80.00%. Importantly, EchoAgent empowers a single system with abilities to learn, observe, operate and reason like an echocardiologist, which holds great promise for reliable Echo interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。