arXiv:2504.13065cs.CV2025-04CVPR被引 18

EchoWorld让超声探头自动导航,精准定位心脏标准切面。

EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance

  • 用运动感知世界模型融合解剖知识与探头移动带来的视觉变化。
  • 在百万级超声图像上训练,引导误差显著低于现有方法。
  • 适合超声辅助诊断、智能医疗设备研发人员使用。

超声心动图对心血管疾病检测至关重要,但依赖经验丰富的超声医师。探头引导系统可提供实时移动指令以获取标准切面图像,是实现AI辅助或全自主扫描的可行方案。然而,构建有效机器学习模型仍具挑战,需理解心脏解剖结构及探头运动与视觉信号间的复杂关系。为此,我们提出EchoWorld——一种面向探头引导的运动感知世界建模框架,能编码解剖知识和运动引起的视觉动态,并有效利用历史视觉-运动序列提升引导精度。该框架采用受世界建模启发的预训练策略,通过预测被掩码的解剖区域并模拟探头调整后的视觉结果。在此预训练模型基础上,引入运动感知注意力机制进行微调,整合历史视觉-运动数据,实现精准自适应引导。模型基于超过200例常规扫描中的一百多万张超声图像训练,定性分析验证其有效捕捉关键超声心动图知识。实验表明,相较现有视觉主干网络与引导框架,本方法在单帧与序列评估协议中均显著降低引导误差。代码已开源:https://github.com/LeapLabTHU/EchoWorld。

原文摘要 · Abstract (English)

Echocardiography is crucial for cardiovascular disease detection but relies heavily on experienced sonographers. Echocardiography probe guidance systems, which provide real-time movement instructions for acquiring standard plane images, offer a promising solution for AI-assisted or fully autonomous scanning. However, developing effective machine learning models for this task remains challenging, as they must grasp heart anatomy and the intricate interplay between probe motion and visual signals. To address this, we present EchoWorld, a motion-aware world modeling framework for probe guidance that encodes anatomical knowledge and motion-induced visual dynamics, while effectively leveraging past visual-motion sequences to enhance guidance precision. EchoWorld employs a pre-training strategy inspired by world modeling principles, where the model predicts masked anatomical regions and simulates the visual outcomes of probe adjustments. Built upon this pre-trained model, we introduce a motion-aware attention mechanism in the fine-tuning stage that effectively integrates historical visual-motion data, enabling precise and adaptive probe guidance. Trained on more than one million ultrasound images from over 200 routine scans, EchoWorld effectively captures key echocardiographic knowledge, as validated by qualitative analysis. Moreover, our method significantly reduces guidance errors compared to existing visual backbones and guidance frameworks, excelling in both single-frame and sequential evaluation protocols. Code is available at https://github.com/LeapLabTHU/EchoWorld.

超声引导世界模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。