arXiv:2512.21691cs.CV2025-12被引 5

揭示VGGT注意力坍缩的动态机制,解释为何长序列会失效

Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective

  • 将全局注意力迭代视为退化扩散过程,建立数学理论模型
  • 发现特征流以1/L速率收敛到狄拉克型分布,预测注意力秩下降规律
  • 解释令牌合并为何能延缓坍缩,为未来3D视觉模型设计提供指导

视觉几何接地变换器(VGGT)在前馈3D重建中达到顶尖性能,但当输入序列超过数百帧时,其全局自注意力层会出现严重坍缩:注意力矩阵迅速退化为近似一秩,令牌几何结构退化至几乎一维子空间,重建误差呈超线性累积。本文从动力学视角出发,将全局注意力迭代建模为退化扩散过程,严格证明在VGGT中,令牌特征流以$O(1/L)$速率收敛至狄拉克型测度,推导出可精确预测经验观察到的秩谱的闭式平均场偏微分方程。该理论与实证注意力热图演化及多项实验结果高度吻合,并解释了其令牌合并修复策略——定期移除冗余令牌——通过降低有效扩散系数来延缓坍缩,且无需额外训练。我们认为该分析为未来可扩展3D视觉变换器提供了原理性视角,并指出其在多模态泛化中的潜力。

原文摘要 · Abstract (English)

Visual Geometry Grounded Transformer (VGGT) delivers state-of-the-art feed-forward 3D reconstruction, yet its global self-attention layer suffers from a drastic collapse phenomenon when the input sequence exceeds a few hundred frames: attention matrices rapidly become near rank-one, token geometry degenerates to an almost one-dimensional subspace, and reconstruction error accumulates super-linearly.In this report,we establish a rigorous mathematical explanation of the collapse by viewing the global-attention iteration as a degenerate diffusion process.We prove that,in VGGT, the token-feature flow converges toward a Dirac-type measure at a $O(1/L)$ rate, where $L$ is the layer index, yielding a closed-form mean-field partial differential equation that precisely predicts the empirically observed rank profile.The theory quantitatively matches the attention-heat-map evolution and a series of experiments outcomes reported in relevant works and explains why its token-merging remedy -- which periodically removes redundant tokens -- slows the effective diffusion coefficient and thereby delays collapse without additional training.We believe the analysis provides a principled lens for interpreting future scalable 3D-vision transformers,and we highlight its potential for multi-modal generalization.

注意力机制3D重建动态系统模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。