arXiv:2608.20284cs.CVcs.RO2026-08

同时预测手术视觉变化与器械轨迹,提升手术规划可靠性。

Towards Surgical World-Action Modeling: A Preliminary Joint Visual-Trajectory Forecasting for Surgical Motion Planning

论文配图:Towards Surgical World-Action Modeling: A Preliminary Joint Visual-Trajectory Forecasting for Surgical Motion Planning
图 1 · 摘自论文原文
  • 联合建模视觉状态与器械运动,实现端到端同步预测
  • 15步预测中,PSNR提升至23.11 dB,ADE降低至22.22像素
  • 适合需要高精度手术动态模拟的研究者与临床辅助系统开发者

可靠的手术规划需同时预判器械运动与术野视觉状态的协同演化。现有方法多将未来场景生成与器械轨迹预测分开处理:仅生成场景的模型无法在轨迹层面评估准确性,仅预测轨迹的模型则忽略运动带来的视觉影响,导致两者一致性缺失。为此,本文提出首个联合视觉-轨迹世界-动作模型,从历史手术观测中同步预测未来视觉状态与器械轨迹。通过时空编码器将视频帧与工具轨迹映射为隐向量,再分别解码生成视觉与轨迹输出。采用分块自回归滚动策略预测15个未来时间步,相较一次性预测显著提升性能:首段PSNR从18.86 dB提升至23.11 dB,平均位移误差(ADE)从45.77像素降至22.22像素。结果验证了联合建模的可行性,但长期预测仍存在视觉退化与轨迹误差累积问题,是未来研究的关键挑战。

原文摘要 · Abstract (English)

Reliable surgical planning requires models to anticipate not only how instruments will move, but also how the operative visual state will evolve together with such motion. Existing approaches typically treat future scene generation and instrument trajectory prediction as two separate tasks. Scene-only models cannot directly evaluate the accuracy of future instrument motion at the trajectory level, while trajectory-only models fail to capture the visual consequences of instrument movement, leaving the consistency between predicted trajectories and future scene evolution unaddressed. Jointly forecasting both provides a more complete account of surgical action-scene dynamics by enabling explicit trajectory-level evaluation while simultaneously modeling the corresponding visual evolution. To bridge this gap, we present a preliminary joint visual-trajectory world-action model that simultaneously forecasts future visual states and instrument trajectories from historical surgical observations. Specifically, we encode historical video frames and tool trajectories into latent representations, which are processed by a temporal-spatial encoder and subsequently decoded through separate visual-state and trajectory prediction heads. Based on this preliminary architecture, a chunked autoregressive rollout is repeatedly applied to predict fifteen future steps. The chunked strategy consistently outperforms direct one-shot prediction across all evaluated horizons, improving first-segment PSNR from 18.86 to 23.11 dB and reducing ADE from 45.77 to 22.22 pixels. These results demonstrate the initial feasibility of joint visual-motion forecasting. However, we observe progressive visual degradation and accumulated trajectory errors over longer prediction horizons, which remain important challenges for future surgical world-action modeling.

手术模拟联合预测轨迹建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。