提出分层潜变量动态建模,有效平衡视频表征的外观与运动能力。
Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives
- 将潜空间拆分为外观和动态子空间,分别加权预测误差
- 混合数据训练下,图像识别提升5.92%,时序推理提升3.21%
- 在细粒度动作任务上性能几乎无损,适合多任务视频学习
联合嵌入预测架构(JEPA)是自监督视频表征学习的有力框架,但小规模训练中辅助目标的行为尚不明确。我们对18种视频JEPA的辅助目标变体进行了小规模实证研究,覆盖单数据集(UCF-101)与混合数据集(UCF-101 + Something-Something V2 + ImageNet-100)两种预训练场景。在三个互补下游基准上评估冻结表征:Diving-48(细粒度运动)、SomethingSomething V2(时间推理)和ImageNet-100(外观)。实验表明,许多辅助目标存在能力权衡:某项下游性能提升常伴随另一项下降。随后我们研究了FWM-HW-LD(带硬区域加权的分层世界模型潜变量动态),该方法将潜变量分解为外观与动态子空间,并对预测误差与潜变量动态误差分别施加硬区域加权。在混合数据设置下,相比基线,其在ImageNet-100上提升+5.92,在SSv2上提升+3.21,而Diving-48性能仅下降0.30个百分点以内。结果表明,潜变量因子化是探索视频JEPA中辅助目标权衡的有效方向。
原文摘要 · Abstract (English)
Joint-Embedding Predictive Architectures (JEPA) are a promising framework for self-supervised video representation learning, yet the behavior of auxiliary objectives in small-scale Video-JEPA training is not well characterized. We report a small-scale empirical study of 18 auxiliary objective variants for Video-JEPA across two pretraining regimes: single-dataset (UCF-101) and mixed-dataset (UCF-101 + Something-Something V2 + ImageNet-100). We evaluate frozen representations on three complementary benchmarks: Diving-48 (fine-grained motion), SomethingSomething V2 (temporal reasoning), and ImageNet-100 (appearance). Our experiments suggest that many auxiliary objectives exhibit capacity trade-offs: gains on one downstream capability often coincide with degradation on another. We then study FWM-HW-LD (Factorized World-Model with Hard-Region-Weighted Latent Dynamics), a training-time objective that separates the latent representation into appearance and dynamics subspaces and applies hard-region weighting to both JEPA prediction errors and latent dynamics errors. In our mixed-dataset setting, FWM-HW-LD improves ImageNet-100 by +5.92 and SSv2 by +3.21 percentage points relative to the reference baseline, while remaining within 0.30 percentage points on Diving-48. These results indicate that latent factorization is a useful direction for studying auxiliary-objective trade-offs in Video-JEPA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。