用视频自蒸馏让单图编码器学会时空感知,提升视觉理解真实性。
Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
- 通过预测下一帧表征,让单图模型学习视频时序规律。
- 仅用2小时视频预训练,ADE20K mIoU从35.0提至36.4。
- 无需光流或追踪,适配现有图像流水线,适合物理感知研究。
自监督图像编码器如DINO近年受到广泛关注,能无标签学习鲁棒视觉特征。但多数方法基于静态图像训练,忽略视频中的时序线索。本文提出一种视频自蒸馏单图编码器,通过从当前帧预测下一帧表示来注入三维空间与时间先验,无需光流或目标追踪。在仅一个2小时视频上预训练后,该方法将ADE20K数据集上的平均交并比(mIoU)从35.0(DoRA)提升至36.4,且可直接替换现有图像流水线。结果表明,视频自蒸馏是一种轻量级的几何感知路径,对构建物理可信世界模型与物理智能至关重要。
原文摘要 · Abstract (English)
Self-supervised image encoders such as DINO have recently gained significant interest for learning robust visual features without labels. However, most SSL methods train on static images and miss the temporal cues inherent in videos. We introduce a video-distilled single-image encoder trained to predict the next-frame representation from the current frame. This simple objective injects 3D spatial and temporal priors without optical flow or tracking. When pre-training on a single 2-hour video, our approach raises the mean Intersection-over-Union (mIoU) on ADE20K from 35.0 (DoRA) to 36.4 while remaining a drop-in replacement for image-only pipelines. Our results highlight video self-distillation as a lightweight route to geometry-aware perception an essential ingredient for physically plausible world models and Physical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。