用动作引导的视觉微调,让手术交互识别更准。
LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

- 用逆动力学模型捕捉动作引发的视觉变化,引导特征学习区域。
- 在多个数据集上提升识别准确率,且特征定位更精准。
- 无需标注框,适合缺乏精细标注的手术视觉任务。
理解器械-组织交互对上下文感知的手术AI和自主机器人手术至关重要。预训练视觉语言模型(VLMs)和视觉编码器通过迁移广泛视觉与语义知识,为传统交互分类器提供替代方案。然而,将其适配到细粒度手术交互仍具挑战:(1) 冻结视觉编码器完全依赖预训练表示,可能保留噪声并提供弱空间定位;(2) 完全微调虽能提升全局语义对齐,但无法确保编码器在正确动作区域学习有意义特征。为此,我们提出LAViFiT,一种端到端的隐式动作引导视觉语言微调框架。逆动力学模型捕捉每种动作引起的视觉变化,前向世界模型驱动编码器表征动作相关区域。片级SIG正则化进一步防止局部特征坍塌,无需额外监督(如边界框或伪标签)。多编码器、多数据集实验表明,该方法显著提升识别性能与图文对齐效果,表征分析显示对完整器械-组织交互区域的锚定更强,空间一致性更高。
原文摘要 · Abstract (English)
Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders offer an alternative to conventional interaction classifiers by transferring broad visual and semantic knowledge. However, adapting them to fine-grained surgical interactions remains challenging: (1) freezing the vision encoder depends entirely on pretrained representations that may retain noise and provide weak spatial localization, while (2) full fine-tuning can improve global semantic alignment without ensuring that the encoder learns meaningful features in the correct action region. We address these limitations by introducing LAViFiT, an end-to-end latent-action-guided framework for vision-language fine-tuning. An inverse dynamics model captures the visual changes induced by each action, while a forward world model drives the encoder to represent action-relevant regions. A patch-level SIG Regularizer further prevents local feature collapse without additional supervision, such as bounding boxes or pseudo-labels. Experiments across multiple encoders and datasets improve recognition and image-text alignment, while representation analyses show stronger grounding over the complete instrument-tissue interaction region and more spatially coherent features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。