arXiv:2509.23672cs.CV2025-09

提出时空信息挖掘的令牌合并方法,提升手术视频理解效率。

Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding

  • 分时空间独立合并冗余令牌,保留关键序列与动态信息
  • 训练免调,减少65%以上计算量,保持竞争力准确率
  • 适合长序列手术视频高效建模,解决计算瓶颈

视觉变压器在手术视频理解任务中表现出色,依赖于对长程依赖关系的建模。然而,现有方法因处理大量时空令牌而面临高昂的计算成本。尽管已有令牌合并工作提升了模型效率,但未能充分考虑视频数据的固有时空结构,也忽略了信息分布的异质性,导致性能不佳。本文提出一种面向手术视频理解的时空信息挖掘令牌合并方法(STIM-TM),为首个专为此任务设计的方法。该方法采用解耦策略,分别独立地在时间与空间维度上减少令牌冗余。时间模块通过显著性加权合并连续帧中对应的空间令牌,保留关键序列信息并维持连续性;空间模块基于时间稳定性分析优先合并静态令牌,保护包含重要手术信息的动态区域。该方法无需训练,实现超过65%的GFLOPs降低,同时在多种手术视频任务中保持优异准确性。该方法还支持长序列手术视频的高效训练,有效缓解手术应用中的计算瓶颈。

原文摘要 · Abstract (English)

Vision Transformer models have shown impressive effectiveness in the surgical video understanding tasks through long-range dependency modeling. However, current methods suffer from prohibitive computational costs due to processing massive spatiotemporal tokens across video frames. While prior work on token merging has advanced model efficiency, they fail to adequately consider the inherent spatiotemporal structure of video data and overlook the heterogeneous nature of information distribution, leading to suboptimal performance. In this paper, we propose a spatiotemporal information mining token merging (STIM-TM) method, representing the first dedicated approach for surgical video understanding. STIM-TM introduces a decoupled strategy that reduces token redundancy along temporal and spatial dimensions independently. Specifically, the temporal component merges spatially corresponding tokens from consecutive frames using saliency weighting, preserving critical sequential information and maintaining continuity. Meanwhile, the spatial component prioritizes merging static tokens through temporal stability analysis, protecting dynamic regions containing essential surgical information. Operating in a training-free manner, STIM-TM achieves significant efficiency gains with over $65\%$ GFLOPs reduction while preserving competitive accuracy across comprehensive surgical video tasks. Our method also supports efficient training of long-sequence surgical videos, addressing computational bottlenecks in surgical applications.

视频理解令牌合并手术视频Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。