arXiv:2510.06809cs.CV2025-10中稿 · Paper被引 1

用视觉动作适配器让超声大模型学会个体心脏结构,提升探头引导精度。

VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance

  • 在超声大模型中嵌入视觉-动作适配器,实时理解个体心脏3D结构
  • 在131万样本数据上表现优于强基线模型,参数量仅为33分之一
  • 适合需要高精度探头引导的临床超声场景,降低操作门槛

超声心动图是检测心脏病的关键工具,但操作难度高导致专业人员短缺。探头引导系统能帮助获取高质量图像,降低操作门槛,但因个体差异大而难以实现稳定性能。这种差异体现在二维图像的低层特征变化和三维解剖结构差异,分别影响图像理解与精准导航。为此,我们利用超声基础模型从海量数据中学到的鲁棒图像表征,并设计视觉-动作适配器(VA-Adapter),在线注入对个体3D结构的理解能力。通过将VA-Adapter嵌入基础模型的图像编码器,模型可从历史视觉-动作序列中推断心脏解剖结构,模拟超声医师的认知过程。在包含超过131万样本的数据集上,实验表明该方法显著优于现有强基线模型,且训练参数量仅需其约1/33。代码已公开于https://github.com/LeapLabTHU/VA-Adapter。

原文摘要 · Abstract (English)

Echocardiography is a critical tool for detecting heart diseases, yet its steep operational difficulty causes a shortage of skilled personnel. Probe guidance systems, which assist in acquiring high-quality images, offer a promising solution to lower this operational barrier. However, robust probe guidance remains challenging due to significant individual variability. This variability manifests as differences in low-level features within two-dimensional (2D) images, which complicates image feature understanding, and differences in individual three-dimensional (3D) structures, which poses challenges for precise navigation. To address these challenges, we first propose leveraging the robust image representations learned by ultrasound foundation models from vast datasets. Yet, applying these models to probe navigation is non-trivial due to their lack of understanding of individual 3D structures. To this end, we meticulously design a Vision-Action Adapter (VA-Adapter) to online inject the capability of understanding individual 3D structures. Specifically, by embedding the VA-Adapter into the foundation model's image encoder, the model can infer cardiac anatomy from historical vision-action sequences, mimicking the cognitive process of a sonographer. Extensive experiments on a dataset with over 1.31M samples demonstrate that the VA-Adapter outperforms strong probe guidance models while requiring approximately 33 times fewer trained parameters. Code is available at https://github.com/LeapLabTHU/VA-Adapter.

超声引导视觉动作适配基础模型心脏成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。