arXiv:2511.14751cs.CVcs.RO2025-11被引 5

用置信度指导视觉几何变压器的令牌合并,实现无重训练加速。

Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers

  • 基于置信度排序并合并低置信度令牌,减少计算量。
  • 在VGGT和Pi3上分别实现21.5倍和20.4倍加速,性能不降。
  • 适用于多视角与流式视觉几何模型,适合实时3D感知场景。

我们提出置信度引导的令牌合并(Co-Me),一种无需重训练或微调即可加速视觉几何变压器的机制。Co-Me通过轻量级置信度预测器对令牌按不确定性排序,选择性合并低置信度令牌,在保持空间覆盖的同时有效降低计算开销。相比基于相似性的合并或剪枝,Co-Me中的置信度信号能可靠指示变压器关注的区域,实现显著加速且不损失性能。该方法可无缝应用于多种多视角及流式视觉几何变压器,加速效果随序列长度增加而提升。在VGGT和Pi3上,分别实现最高21.5倍和20.4倍加速,使视觉几何变压器可用于实时3D感知与重建。

原文摘要 · Abstract (English)

We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the base model. Co-Me distilled a light-weight confidence predictor to rank tokens by uncertainty and selectively merge low-confidence ones, effectively reducing computation while maintaining spatial coverage. Compared to similarity-based merging or pruning, the confidence signal in Co-Me reliably indicates regions emphasized by the transformer, enabling substantial acceleration without degrading performance. Co-Me applies seamlessly to various multi-view and streaming visual geometric transformers, achieving speedups that scale with sequence length. When applied to VGGT and Pi3, Co-Me achieves up to 21.5x and 20.4x speedup, making visual geometric transformers practical for real-time 3D perception and reconstruction.

视觉几何模型加速令牌合并

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。