arXiv:2512.04085cs.CV2025-12被引 1

用一个人的视角视频训练模型,发现不同人学出的视觉理解高度一致。

Unique Lives, Shared World: Learning from Single-Life Videos

  • 只用单个人的视角视频自监督训练视觉编码器,利用多视角信息学习
  • 单人生视频训练30小时,性能媲美30小时网络数据,跨环境可迁移
  • 适合研究个体经验如何塑造通用视觉理解的学者

我们提出「单人生」学习范式,即仅用一个人的内省视角视频训练独立视觉模型。利用同一人生中自然产生的多视角信息,实现自监督视觉编码器学习。实验表明:第一,不同人生训练的模型在几何理解上高度对齐,通过新提出的交叉注意力度量验证了模型内部表征的功能一致性;第二,单人生模型学到的几何表征具有强泛化能力,可在未见环境中有效迁移至深度估计等下游任务;第三,仅用一周中30小时同一个人的视频训练,性能即可媲美30小时多样化的网络视频数据,凸显单人生表征学习的强大潜力。整体结果表明,世界的共享结构不仅使个体生命训练的模型趋于一致,更提供了强大的视觉表征学习信号。

原文摘要 · Abstract (English)

We introduce the "single-life" learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual. We leverage the multiple viewpoints naturally captured within a single life to learn a visual encoder in a self-supervised manner. Our experiments demonstrate three key findings. First, models trained independently on different lives develop a highly aligned geometric understanding. We demonstrate this by training visual encoders on distinct datasets each capturing a different life, both indoors and outdoors, as well as introducing a novel cross-attention-based metric to quantify the functional alignment of the internal representations developed by different models. Second, we show that single-life models learn generalizable geometric representations that effectively transfer to downstream tasks, such as depth estimation, in unseen environments. Third, we demonstrate that training on up to 30 hours from one week of the same person's life leads to comparable performance to training on 30 hours of diverse web data, highlighting the strength of single-life representation learning. Overall, our results establish that the shared structure of the world, both leads to consistency in models trained on individual lives, and provides a powerful signal for visual representation learning.

自监督学习视觉表征内省视频几何理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。