arXiv:2603.13741cs.CV2026-03被引 2

1000段多视角第一人称视频,助力智能眼镜场景重建研究

Ego-1K -- A Large-Scale Multiview Video Dataset for Egocentric Vision

  • 12摄像头同步采集第一人称动态视频,聚焦手部与物体交互
  • 真实场景中存在大视差和运动模糊,挑战现有3D/4D重建方法
  • 适合研究智能眼镜视觉、动态场景建模的团队使用

我们提出Ego-1K,一个大规模的时间同步第一人称多视角视频数据集,用于推动神经3D视频合成与动态场景理解。数据集包含近1000段短视频,由环绕4摄像头VR头显的定制12相机阵列采集,记录用户佩戴状态下的自然动作。内容聚焦不同场景中的手部运动与手物交互。本文详述了采集设备设计、数据处理流程与标定方法。该数据集为第一人称场景重建提供了新基准,这一方向在多摄像头智能眼镜普及的背景下日益重要。实验表明,由于近距离动态物体与设备自身运动带来的大视差和图像运动,现有3D与4D新视角合成方法面临独特挑战。数据集已开放下载:https://huggingface.co/datasets/facebook/ego-1k。

原文摘要 · Abstract (English)

We present Ego-1K, a large-scale collection of time-synchronized egocentric multiview videos designed to advance neural 3D video synthesis and dynamic scene understanding. The dataset contains nearly 1,000 short egocentric videos captured with a custom rig with 12 synchronized cameras surrounding a 4-camera VR headset worn by the user. Scene content focuses on hand motions and hand-object interactions in different settings. We describe rig design, data processing, and calibration. Our dataset enables new ways to benchmark egocentric scene reconstruction methods, an important research area as smart glasses with multiple cameras become omnipresent. Our experiments demonstrate that our dataset presents unique challenges for existing 3D and 4D novel view synthesis methods due to large disparities and image motion caused by close dynamic objects and rig egomotion. Our dataset supports future research in this challenging domain. It is available at https://huggingface.co/datasets/facebook/ego-1k.

第一人称视觉多视角视频场景重建智能眼镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。