从第一视角视频中无监督学习稳定物体表征,提升视觉理解能力。
Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video
- 通过帧内蒸馏、深度约束和时序一致性三机制联合发现并稳定物体原型。
- 在无监督物体发现上提升8.0%的CorLoc,语义分割提升4.8% mIoU。
- 适合构建具身智能系统的视觉基础,尤其适用于真实场景视频理解。
人类通过感知与环境互动发展视觉智能——一种基于第一人称经验的自监督学习过程。受此启发,我们探讨如何让人工系统从连续、未经筛选的第一视角视频中,无需人工标注即可学习稳定的物体表征。这一任务面临杂乱、遮挡和自身运动带来的分离、识别与持续跟踪挑战。我们提出EgoViT,一种统一的视觉Transformer框架,用于从无标签第一视角视频中学习稳定物体表征。EgoViT通过三种协同机制启动学习:(1) 原型物体学习,利用帧内蒸馏生成判别性表示;(2) 深度正则化,将表示锚定于几何结构;(3) 教师过滤的时序一致性,保证身份恒定。这形成良性循环,使初始物体假设逐步优化为稳定持久的表征。该框架在无标签第一人称视频上端到端训练,对不同来源与质量的几何先验具有鲁棒性。在标准基准上,EgoViT在无监督物体发现中实现+8.0% CorLoc提升,在语义分割中实现+4.8% mIoU提升,展现出为具身智能提供稳健视觉抽象的基础潜力。
原文摘要 · Abstract (English)
Humans develop visual intelligence through perceiving and interacting with their environment - a self-supervised learning process grounded in egocentric experience. Inspired by this, we ask how can artificial systems learn stable object representations from continuous, uncurated first-person videos without relying on manual annotations. This setting poses challenges of separating, recognizing, and persistently tracking objects amid clutter, occlusion, and ego-motion. We propose EgoViT, a unified vision Transformer framework designed to learn stable object representations from unlabeled egocentric video. EgoViT bootstraps this learning process by jointly discovering and stabilizing "proto-objects" through three synergistic mechanisms: (1) Proto-object Learning, which uses intra-frame distillation to form discriminative representations; (2) Depth Regularization, which grounds these representations in geometric structure; and (3) Teacher-Filtered Temporal Consistency, which enforces identity over time. This creates a virtuous cycle where initial object hypotheses are progressively refined into stable, persistent representations. The framework is trained end-to-end on unlabeled first-person videos and exhibits robustness to geometric priors of varied origin and quality. On standard benchmarks, EgoViT achieves +8.0% CorLoc improvement in unsupervised object discovery and +4.8% mIoU improvement in semantic segmentation, demonstrating its potential to lay a foundation for robust visual abstraction in embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。