arXiv:2602.20500cs.ROcs.CV2026-02

用事件图挖掘手术相机控制策略,提升自动化精度与医生可理解性。

Strategy-Supervised Autonomous Laparoscopic Camera Control via Event-Driven Graph Mining

  • 通过事件图结构化手术视频,提取可复用的相机操作策略
  • 在离体实验中降低视场中心误差35.26%、图像抖动62.33%
  • 支持语音输入的人机协同,适合手术机器人系统研发者

自主腹腔镜摄像机控制需在快速器械-组织交互中保持稳定安全视野,同时对医生可解释。本文提出一种基于策略的框架,结合高层视觉-语言推理与底层闭环控制。离线阶段,将原始手术视频解析为与相机相关的时序事件(如操作、工作距离偏移、视野质量下降),构建属性事件图;挖掘这些图得到一组紧凑的可复用相机操作策略原型,用于监督学习。在线阶段,微调后的视觉-语言模型(VLM)处理实时腹腔镜画面,预测主导策略及基于图像的离散运动指令,由满足严格安全约束的IBVS-RCM控制器执行;支持语音输入实现直观人机协同。在医生标注数据集上,事件解析实现可靠的时间定位(F1得分0.86),挖掘策略与专家理解具有强语义一致性(聚类纯度0.81)。大量离体实验在硅胶模型和猪组织上验证,系统性能优于初级外科医生,在标准化相机操控评估中,视场中心误差减少35.26%,图像抖动降低62.33%,同时保持平稳运动与稳定工作距离调节。

原文摘要 · Abstract (English)

Autonomous laparoscopic camera control must maintain a stable and safe surgical view under rapid tool-tissue interactions while remaining interpretable to surgeons. We present a strategy-grounded framework that couples high-level vision-language inference with low-level closed-loop control. Offline, raw surgical videos are parsed into camera-relevant temporal events (e.g., interaction, working-distance deviation, and view-quality degradation) and structured as attributed event graphs. Mining these graphs yields a compact set of reusable camera-handling strategy primitives, which provide structured supervision for learning. Online, a fine-tuned Vision-Language Model (VLM) processes the live laparoscopic view to predict the dominant strategy and discrete image-based motion commands, executed by an IBVS-RCM controller under strict safety constraints; optional speech input enables intuitive human-in-the-loop conditioning. On a surgeon-annotated dataset, event parsing achieves reliable temporal localization (F1-score 0.86), and the mined strategies show strong semantic alignment with expert interpretation (cluster purity 0.81). Extensive ex vivo experiments on silicone phantoms and porcine tissues demonstrate that the proposed system outperforms junior surgeons in standardized camera-handling evaluations, reducing field-of-view centering error by 35.26% and image shaking by 62.33%, while preserving smooth motion and stable working-distance regulation.

手术机器人视觉语言模型事件图相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。