arXiv:2608.15009cs.ROcs.CV2026-08

让超声探头懂力感,自动稳定扫描

ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning

论文配图:ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning
图 1 · 摘自论文原文
  • 融合视觉与力觉信号,实时指导探头运动
  • 五种临床视图下,压力控制稳定性提升41%
  • 适合医疗机器人、智能超声系统研发者

具身智能超声扫描通过整合感知、决策与执行能力,实现检查流程的自动化与标准化。然而,现有方法在力觉与超声模态间耦合松散,且缺乏对扫描阶段的感知,难以捕捉动态探头-组织交互。为此,我们提出ForceU-VLA,一种面向自主具身超声扫描的力觉感知视觉-语言-动作模型,全程利用力信号与超声图像反馈,实现高质量、高稳定的超声采集。首先,提出力-超声协同融合模块(FUSFM),协同融合视觉与力觉信息,为探头运动提供稳定可靠引导;其次,设计阶段自适应调制机制(SAMM),根据不同扫描阶段动态调节多模态特征,提升表征质量。此外,构建了ForceU-VLA-Data,一个真实世界、含力觉信号的具身超声数据集,涵盖两器官、五种典型临床视图,包含450条专家采集轨迹,约10万帧同步多模态数据。大量实验表明,ForceU-VLA显著提升了接触稳定性与探头压力调控能力,有效增强任务执行质量与系统可靠性。源代码已公开于https://github.com/VMVLab/ForceU-VLA。

原文摘要 · Abstract (English)

Embodied intelligent ultrasound scanning enables the automation and standardization of the ultrasound examination process by integrating perception, decision-making, and execution capabilities. However, existing methods suffer from loosely coupled modeling between force and ultrasound modalities and lack awareness of scanning stages, which limits their ability to capture dynamic probe-tissue interactions. To address these issues, we propose ForceU-VLA, a force-aware Vision-Language-Action model for autonomous embodied ultrasound scanning, which leverages force signals and ultrasound image feedback throughout the scanning process to enable accurate and high-quality ultrasound acquisition. Firstly, we propose a Force-Ultrasound Synergistic Fusion Module (FUSFM) that synergistically fuses ultrasound visual and force-feedback information to provide stable, reliable guidance for probe motion. Secondly, a Stage-Adaptive Modulation Mechanism (SAMM) is proposed to accommodate the task requirements across different scanning stages by adaptively modulating multimodal features to enhance their representation quality. Additionally, we introduce ForceU-VLA-Data, a real-world, force-aware embodied ultrasound dataset that integrates visual, force, and action signals, including data from two organs across five representative clinical scanning views, and comprising 450 expert-collected trajectories with approximately 100,000 synchronized multimodal frames. Extensive experimental results demonstrate that ForceU-VLA significantly improves contact stability and probe pressure regulation in embodied ultrasound scanning, thereby effectively enhancing task execution quality and overall system reliability. The source code is available at https://github.com/VMVLab/ForceU-VLA.

具身智能超声扫描力觉感知多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。