arXiv:2608.02471cs.CVcs.AI2026-08

用手术动作反推组织可操作性,实现自动镜头追踪,降低医生认知负担。

Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

  • 通过动作轨迹逆向生成密集标注,替代人工标注组织交互点。
  • 自动构图系统在12例胆囊切除术中显著减少医生对助手的口头指令。
  • 模型跨手术类型迁移有效,方向一致性达95.16%,适合智能手术辅助场景。

在腹腔镜手术中,医生视线会跟随器械将要作用的位置;通过视觉注意力建模减轻这一负担,需要密集标注这些交互位置。这些标注蕴含隐性知识:专家对关键位置有共识,但难以表述规则。本文展示可通过完成的手术视频恢复此类标注,将记录的器械轨迹转化为密集连续的监督信号。DiffeoAfford通过将器械尖端附着于组织并使用微分同胚约束跟踪进行形变传输,匹配上下文知情标注者的准确性。该模型仅基于这些标签训练,未使用注视数据,却在空间和时间上比摄像头助手更贴近医生视线。框架还可跨手术迁移:在子宫切除术视频上,独立训练的预测器与后续相机运动方向一致性达95.16%。在12对胆囊切除术(共24例)中,自动构图应用AffordView能提前将预测目标置于画面中心,通过主观、生理和行为指标综合证明降低了医生的认知负荷,包括减少对摄像助手的口头指令次数。从动作中推导监督信号,为前瞻式辅助提供了可扩展路径。

原文摘要 · Abstract (English)

In laparoscopy, surgeon gaze tracks where the instruments will act; easing this demand through visual attention modeling requires dense labels of those interaction loci. These encode tacit knowledge: experts converge on consensus loci yet struggle to state the rules. Here we show that such labels can be recovered from completed actions in surgical videos, in which recorded instrument trajectories are converted into dense, continuous supervision. DiffeoAfford grounds tissue affordance by attaching instrument tips to the tissue and transporting them through deformation using diffeomorphism-constrained tracking, matching context-informed annotators' accuracy. Trained on these labels and never on gaze, a real-time model aligns with surgeon gaze more closely in space and time than does camera-assistant gaze. The framework also transfers across procedures: on hysterectomy videos, a separately trained predictor reaches 95.16% directional consistency with subsequent camera motion. In 12 paired cholecystectomies (24 procedures), the auto-framing application AffordView, which proactively centers predicted targets in view, lowered surgeon cognitive workload on converging subjective, physiological, and behavioral measures, including a reduced number of verbal instructions to the camera assistant. Deriving supervision from action rather than manual annotation offers a scalable route to anticipatory assistance.

手术辅助自动构图认知负荷动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。