用学习的超声动态模型实现机器人探头自动导航,提升颈部扫描成功率。
Action-Conditioned World Model for Goal Plane Probe Guidance in Robotic Ultrasound

- 构建动作条件世界模型,预测探头运动下的超声图像变化。
- 在真实场景中实现颈动脉导航成功率70.0%,甲状腺导航65.0%。
- 适合研究机器人超声自主导航与医疗影像生成的团队参考。
我们提出一种面向机器人超声颈部扫描的目标平面探头引导的动作条件世界模型框架。自主超声任务通常需要大量探头运动轨迹训练,但高质量示范数据收集成本高,且因接触力、组织形变和视角依赖的声学伪影,难以构建精确仿真器。为此,我们采用两阶段基于模型的学习流程:首先,一个潜在条件扩散世界模型从近期图像帧、探头运动及时间偏移量预测未来超声观测;其次,一个目标条件时序变压器预测有序探头运动,并通过冻结的世界模型提供的奖励进行微调。在自采集数据集上的实验表明,该世界模型能保留目标导向扫描中的动作相关解剖结构。在真实闭环实验中,系统在颈动脉引导任务中达到70.0%的成功率,在甲状腺引导任务中达到65.0%。结果验证了学习到的超声动力学在训练目标导向机器人探头导航中的潜力。
原文摘要 · Abstract (English)
We present an action-conditioned world model framework for goal plane probe guidance in robotic ultrasound, with a focus on neck ultrasound scanning. Autonomous ultrasound tasks often require large numbers of probe-motion trajectories for training, but collecting high-quality demonstrations is labor-intensive and explicit simulators are difficult to build because ultrasound appearance depends on contact, tissue deformation, and view-dependent acoustic artifacts. We address this problem with a two-stage model-based learning pipeline. First, a latent conditional diffusion world model predicts future ultrasound observations from recent context frames, probe motions and temporal offset. Second, a goal-conditioned temporal transformer predicts ordered probe motions and is fine-tuned using rewards from the frozen world model. Experiments on the self-collected dataset show that the world model preserves action-dependent anatomical structure on target-directed scans. In real-world closed loop experiments, the framework achieves success rates of 70.0\% for carotid guidance and 65.0\% for thyroid guidance. These results demonstrate the potential of learned ultrasound dynamics for training goal-directed robotic probe navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。