arXiv:2603.06999cs.CV2026-03

通过轨迹条件嵌入提升手术器械-组织交互识别精度

TrajPred: Trajectory-Conditioned Joint Embedding Prediction for Surgical Instrument-Tissue Interaction Recognition in Vision-Language Models

  • 用器械运动轨迹编码时间信息,生成更精准的视觉语义嵌入
  • 在CholecT50数据集上提升平均精度和前K准确率
  • 适合需要细粒度动作理解的手术视觉语言模型研究

识别手术器械与组织的交互对构建情境感知的智能助手至关重要。视觉语言模型(VLMs)为手术感知提供了新路径,在多种任务上表现出优于传统专用深度学习方法的泛化能力。然而,其在器械-组织交互识别上的表现仍受限,主要源于两点:(1) 多数模型未能有效利用时间信息;(2) 视觉与文本间的对齐常遗漏细粒度动作细节。为此,我们提出TrajPred框架,通过编码器械轨迹引入时间运动线索,并在轨迹条件下引入预测模块,生成更能捕捉细粒度动作细节的视觉语义嵌入。此外,结合提示调优和动词重述技术,实现对器械-组织交互识别任务的顺畅适配。在公开的腹腔镜基准数据集CholecT50上的大量实验表明,该方法提升了平均精度和Top-K准确率。我们还通过可视化对比视觉与文本嵌入之间的余弦相似度,探究了交互区域视觉嵌入与对应文本的对齐情况。结果表明,所提方法增强了相关视觉与文本表示间的对齐效果。

原文摘要 · Abstract (English)

Recognizing instruments' interactions with tissues is essential for building context-aware AI assistants in robotic surgery. Vision-language models (VLMs) have opened a new avenue for surgical perception and achieved better generalization on a wide range of tasks compared to conventional task-specific deep learning approaches. However, their performance on instrument--tissue interaction recognition remains limited, largely due to two challenges: (1) many models do not effectively leverage temporal information, and (2) alignment between vision and text often misses fine-grained action details. To address these issues, we propose TrajPred, a framework that encodes instrument trajectories to incorporate temporal motion cues and, conditioned on these trajectories, introduces a predictor module to generate visual semantic embeddings that better capture fine-grained action details. We further incorporate prompt tuning and a verb-rephrasing technique to enable smooth adaptation to the instrument--tissue interaction recognition task. Extensive experiments on the public laparoscopic benchmark, CholecT50, show that our method improves both Average Precision and Top-K accuracy. We also investigate whether visual embeddings of instrument--tissue interaction regions align better with the corresponding text by visualizing the cosine similarity between visual and textual embeddings. The visualization results indicate that the proposed method improves alignment between relevant visual and textual representations.

手术视觉视觉语言轨迹建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。