arXiv:2601.09040cs.CV2026-01

不靠端到端反向传播,也能训练视频ViT,效果接近传统方法。

Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers

  • 将编码器分块,每块用局部重建损失独立优化。
  • 在多种模型规模下,性能接近端到端基线,线性探测与检索任务表现良好。
  • 揭示了早期结构暴露快、后期块饱和的表征发展规律,适合研究学习动态者。

端到端反向传播通过全局误差信号耦合所有层,实现协同学习但需长程信用分配。受块级自监督学习(BWSSL)进展启发,我们探讨掩码视频变压器是否可在无需端到端反向传播的情况下训练。将BWSSL应用于掩码视频建模仍相对未被充分探索,且需处理时空上下文和长程时间结构。更广泛地,关于BWSSL与端到端训练在学习动态和深度表征发展方面的比较分析仍较少。我们通过将编码器划分为若干块,并对每块使用局部掩码重建损失进行优化,将块级学习应用于掩码自编码视频视觉变换器。在不同模型尺寸和划分粒度下,训练均收敛,且在线性探测和检索代理任务上生成的表征接近匹配的端到端基线。为比较中间表征,我们分析了深度层面可解码性、块间相似性及像素级诊断。块级训练更早暴露高层结构,而后期块趋于饱和,并进入更具几何保真性的运行状态。它还可能引发与更强早期混合一致的令牌级变化,而这些变化是聚合指标所忽略的。这些发现表明,后期块饱和和接口形成是剩余差距的贡献因素。

原文摘要 · Abstract (English)

End-to-end backpropagation couples all layers through a global error signal, enabling coordinated learning but requiring long-range credit assignment. Motivated by recent progress in blockwise self-supervised learning (BWSSL), we ask whether masked video transformers can be trained without end-to-end backpropagation. Applying BWSSL to masked video modeling remains relatively underexplored and must handle spatiotemporal context and long-range temporal structure. More broadly, analyses that compare BWSSL and end-to-end training in terms of learning dynamics and depth-wise representation development remain sparse. We apply blockwise learning to a masked autoencoding video vision transformer by partitioning the encoder into blocks, each of which is optimized with a local masked reconstruction loss. Across model sizes and partition granularities, training converges and yields representations close to matched end-to-end baselines under linear-probe and retrieval proxies. In order to compare intermediate representations, we analyze depth-wise decodability, inter-block similarity, and patch-level diagnostics. Blockwise training exposes higher-level structure earlier, while later blocks saturate and operate in a more geometry-preserving regime. It can also induce token-level shifts consistent with stronger early mixing that pooled metrics can miss. These findings point to late-block saturation and interface formation as contributors to the remaining gap.

视频视觉自监督块级学习表征发展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。