通过分层压缩提升VGGT效率,实现6.7倍加速且不损失重建质量。
RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

- 按层级差异设计压缩策略,区分浅层、中层和深层的冗余特性。
- 在保持几何结构与姿态关键路径的前提下,实现6.7倍推理加速。
- 无需训练即可部署,适合追求高效3D重建的视觉系统开发者。
视觉几何接地变压器(VGGT)可在一次前向传播中从多视角图像恢复稠密3D场景结构,但其二次复杂度的跨帧注意力限制了可扩展性。现有无训练加速方法沿单一轴线均匀降维,忽略了层级异质性。我们的谱分析、探测分析与因果分析揭示三种状态:浅层缺乏跨视角结构,中层驱动跨视角对齐,深层虽对稠密几何冗余,但其跨帧注意力对姿态仍至关重要。RegimeVGGT采用双轴分层U型压缩:显著性引导带状合并保护几何与边缘显著标记;选择性保护的键值下采样保留跨帧空间覆盖及姿态关键路径,通过相位偏移空间网格、参考帧锚点和未压缩的相机/注册标记实现。该方法无需训练,在保持重建质量前提下,相比VGGT*实现6.7倍加速。
原文摘要 · Abstract (English)
Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7x speedup over VGGT* at matched reconstruction quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。