arXiv:2511.13948cs.CVcs.CL2025-11被引 3

用智能体框架实现超声心动图的自动化测量与解读

EchoAgent: Guideline-Centric Reasoning Agent for Echocardiography Measurement and Interpretation

  • 在大模型控制下调度视觉工具,完成时空定位与测量
  • 预测每帧可测量性,自动选择最优分析工具
  • 结果符合临床指南,适合医疗AI可信化研究

目的:超声心动图解读需要视频级推理和基于指南的测量分析,现有深度学习模型无法支持。我们提出EchoAgent框架,实现该领域的结构化、可解释自动化。方法:在大语言模型(LLM)控制下,协调专用视觉工具完成时序定位、空间测量和临床解读。关键贡献是测量可行性预测模型,判断每帧中解剖结构是否可可靠测量,从而实现自主工具选择。我们构建了一个包含多样、临床验证视频-查询对的基准数据集用于评估。结果:尽管增加了时空分析复杂性,EchoAgent仍能实现准确、可解释的结果,输出基于视觉证据和临床指南,支持透明性和可追溯性。结论:本工作展示了基于任务特异性工具和全视频级自动化的智能体式、指南对齐推理在超声心动图分析中的可行性,为心脏超声领域可信AI开辟新方向。

原文摘要 · Abstract (English)

Purpose: Echocardiographic interpretation requires video-level reasoning and guideline-based measurement analysis, which current deep learning models for cardiac ultrasound do not support. We present EchoAgent, a framework that enables structured, interpretable automation for this domain. Methods: EchoAgent orchestrates specialized vision tools under Large Language Model (LLM) control to perform temporal localization, spatial measurement, and clinical interpretation. A key contribution is a measurement-feasibility prediction model that determines whether anatomical structures are reliably measurable in each frame, enabling autonomous tool selection. We curated a benchmark of diverse, clinically validated video-query pairs for evaluation. Results: EchoAgent achieves accurate, interpretable results despite added complexity of spatiotemporal video analysis. Outputs are grounded in visual evidence and clinical guidelines, supporting transparency and traceability. Conclusion: This work demonstrates the feasibility of agentic, guideline-aligned reasoning for echocardiographic video analysis, enabled by task-specific tools and full video-level automation. EchoAgent sets a new direction for trustworthy AI in cardiac ultrasound.

医学影像智能体超声心动图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。