arXiv:2608.16074cs.ROcs.CV2026-08

让超声扫描自动执行,基于语义指令和实时影像反馈

US-VLA: An Ultrasound Vision-Language-Action Model for Embodied Abdomina

论文配图:US-VLA: An Ultrasound Vision-Language-Action Model for Embodied Abdomina
图 1 · 摘自论文原文
  • 用视觉-语言-动作模型融合超声图像与临床目标
  • 在320段专家扫描轨迹上实现精准探头控制
  • 适合想开发智能超声系统的研究者和医生

人工智能辅助超声扫描通过实时指导标准化成像、降低操作者依赖,提升诊断可靠性与效率。然而,现有强化学习或学习辅助扫描方法通常依赖精心设计的奖励函数或大量交互数据,限制了其在不同设备、患者群体及复杂临床场景下的泛化能力与稳定性。为此,我们提出一种超声视觉-语言-动作模型(US-VLA),显式编码临床语义目标,并在实时超声反馈下生成序列化探头操作动作。首先,设计超声感知的专家融合模块,联合整合超声观测与辅助上下文信息,使语义层面的超声反馈有效引导扫描过程。其次,构建US-VLA-Data真实世界数据集,涵盖肝肾检查,包含五个标准切面,共320条专家扫描轨迹,约8万组同步时间步数据。大量实验表明,US-VLA在超声探头操控任务中表现优异,证明其在评估的腹部超声场景中具有有效性与良好泛化能力。源代码已开源。

原文摘要 · Abstract (English)

Artificial intelligence-assisted ultrasound scanning enhances diagnostic reliability and efficiency by providing real-time guidance for standardized image acquisition and reducing operator dependence. However, existing reinforcement learning and learning-assisted ultrasound scanning methods typically rely on carefully designed reward functions or extensive interaction data, which limits their generalization ability and stability across different devices, patient populations, and complex clinical scenarios. To address these challenges, we propose an ultrasound vision-language-action model (US-VLA) for automated ultrasound scanning that explicitly encodes clinical semantic goals and generates sequential probe manipulation actions under real-time ultrasound feedback. In particular, we first design an ultrasound-aware expert fusion module to jointly integrate ultrasound observations with auxiliary contextual information, enabling semantic ultrasound feedback to effectively guide the scanning process. Then, we construct US-VLA-Data, a real-world dataset covering liver and kidney examinations, which includes five clinically defined standard planes and comprises 320 expert scanning trajectories with approximately 80,000 synchronized timesteps. Extensive experiments demonstrate that US-VLA achieves competitive performance in ultrasound probe manipulation tasks, indicating its effectiveness and promising generalization within the evaluated abdominal ultrasound setting. The source code is available at https://github.com/VMVLab/US-VLA.

超声辅助视觉语言机器人扫描医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。