将手术动作的骨骼运动转化为图像可控信号,实现精准的手术视频生成。
From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

- 用五种图像对齐的控制模态,把骨骼运动映射到视觉空间。
- 通过分层路由机制,动态选择关键控制信号,提升生成效率与精度。
- 适合需要高精度动作控制的机器人手术视频生成研究者。
动作条件下的手术视频生成是机器人手术中的关键挑战,核心难点在于低维控制向量需精确调控复杂的图像空间演化。本文提出一种从骨骼运动到视觉控制的映射范式,将关节运动转化为统一的五种图像对齐控制模态。在此基础上,构建分层路由的视觉控制框架,动态选择最相关的控制模态与运动尺度,避免全量信号叠加。设计基于骨骼先验的路由损失函数,确保生成结果物理合理、时间稳定且专家利用高效。为提升效率,提出预算化训练与推理方案,利用路由带来的稀疏性,在训练和执行中剔除低重要性路径,实现自适应计算,互补于传统蒸馏方法。同时构建新基准数据集,通过人机协同语义标注与可微姿态追踪获取精细关节标注,提供真实监督。大量实验表明,该方法在动作忠实度、视觉保真度及跨域泛化上均优于多种基线。其高效变体显著降低延迟,同时保持强控制精度。
原文摘要 · Abstract (English)
Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we propose a kinematic-to-visual lifting paradigm that converts articulated kinematics into a unified set of five image-aligned control modalities. Building on this representation, we introduce a hierarchically routed visual control framework that selectively activates the most relevant control modalities and motion scales. Instead of uniformly applying all control signals, our model performs hierarchical routing to dynamically allocate conditioning capacity. We further design kinematic-prior-guided routing loss functions to ensure physically meaningful, temporally stable, and efficient expert utilization. To improve efficiency, we propose a budgeted training and inference scheme that leverages routing-induced sparsity. By selectively discarding low-significance control pathways during training and execution, our approach enables adaptive computation that is complementary to standard distillation. We additionally construct a new benchmark with curated articulated annotations, obtained through human-in-the-loop semantic labeling and differentiable pose tracking, providing realistic supervision for action-conditioned surgical video generation. Extensive experiments demonstrate that our method consistently improves action faithfulness, visual fidelity, and cross-domain generalization over diverse baselines. Moreover, our efficient variant achieves substantial reductions in latency while maintaining strong control accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。