提出动态调整手术动作执行范围,提升双臂机器人操作的可靠性。
Trajectory Divergence Horizon Decision for Reliable Dual-Arm Surgical Subtask Manipulation

- 通过轨迹发散度评估动作可靠性,动态决定执行长度
- 实测成功率提升:针操作从55%到60%,组织操作从55%到80%
- 适合需要高可靠性的手术机器人自主控制场景
随着临床工作量增加,手术机器人系统日益普及,推动重复性操作子任务的自动化。基于学习的控制器相比规则和解析方法更具泛化能力,但多数仅针对单一任务训练,难以跨术式复用。视觉-语言-动作(VLA)模型整合视觉感知、语言对齐与动作生成,为更可组合的手术自主提供可能。然而现有VLA策略依赖固定长度开环动作序列,场景变化易导致误差累积,带来风险。为此,本文将VLA部署建模为自适应执行时域决策问题,提出轨迹发散时域决策(TDHD)机制,在测试时通过小扰动下两流匹配轨迹的差异衡量每步动作可靠性,并采用双阈值截断规则触发及时重规划。我们构建了类达芬奇双臂真实场景基准,含同步多视角感知与语言指令,采集600次远程操作示范,涵盖针(抓取、重握)与组织(抬升、切除)操作。在真实硬件上每任务设置20次试验,TDHD显著优于最新VLA基线:针操作成功率从55%提升至60%,组织操作从55%提升至80%,尤其在最终操作阶段增益最大。结果表明,自适应执行控制对可靠部署VLA模型于手术机器人操作至关重要。
原文摘要 · Abstract (English)
Surgical robotic systems are increasingly being adopted as clinical workload rises, motivating autonomous solutions for repetitive manipulation subtasks. Learning-based controllers improve generalization compared with rule-based and analytic approaches, but most are trained for individual tasks and remain difficult to reuse across procedures. Vision-Language-Action (VLA) models provide a unified framework that integrates visual perception, language grounding, and action generation, offering a promising path toward more composable surgical autonomy. However, existing VLA policies rely on fixed-length open-loop action sequences, where changing scene conditions can lead to accumulated errors and potential risks in surgical manipulation. To mitigate this issue, we formulate surgical VLA deployment as an adaptive execution-horizon decision problem and propose Trajectory Divergence Horizon Decision (TDHD), a test-time mechanism that estimates step-wise action reliability by measuring the divergence between two flow-matching-generated trajectories under small noise perturbations and truncates execution using a dual-threshold rule to trigger timely replanning. We further establish a real-world da Vinci-like dual-arm benchmark with synchronized multi-view perception and language instructions, and collect 600 teleoperated demonstrations across needle (reach, pick, regrasp) and tissue (reach, lift, resection) manipulation suites. On real hardware with 20 trials per task setting, TDHD consistently improves performance over the latest VLA baselines: success increases from 55\% to 60\% for needle manipulation and from 55\% to 80\% for tissue manipulation, with the largest gains observed in the final manipulation stages. These results highlight the importance of adaptive execution control for reliable deployment of VLA models in surgical robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。