用手机采集超小时长第一人称数据,突破机器人数据收集硬件门槛
MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware

- 基于手机传感器实现长时间第一人称姿态追踪
- 构建200小时多样数据集,支持复杂任务长期依赖建模
- 开源工具链+免费应用,人人可参与数据采集
视觉-语言-动作(VLA)模型推动对大规模第一人称数据集的需求,但长期轨迹采集的硬件与基础设施仍难以获取。当前数据集通常仅包含数分钟的片段,无法捕捉复杂机器人任务所需的长期时间依赖关系。我们提出MobileEgo Anywhere框架,可在消费级移动设备上采集超过一小时的第一人称轨迹,利用现代智能手机传感器实现长期姿态追踪,突破传统机器人数据采集的硬件限制。我们发布三个组件:(1) STERA,一个开源视频处理流水线,将原始手机录制转换为标准化、可用于VLA与基础模型研究的训练格式;(2) 一款免费移动端应用,允许任意用户记录第一人称活动;(3) 一个包含584个会话、总时长200小时的多样化长序列第一人称数据集,支持跨会话的状态持续追踪。进一步实验表明,该数据集可作为有效训练信号:在训练中使用该数据能降低后续动作预测的误差。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have driven demand for large-scale egocentric datasets, yet the hardware and infrastructure to collect long-horizon data remain inaccessible. Datasets today typically have episodes only a few minutes long, which fails to capture the long-horizon temporal dependencies that complex robotic task execution requires. We present MobileEgo Anywhere, a framework for collecting hour-plus egocentric trajectories on commodity mobile hardware that uses modern smartphone sensors for long-term pose tracking without the hardware barriers of traditional robotics data collection. We release three components: (1) STERA, an open-source video-processing pipeline that converts raw mobile captures into standardized, training-ready formats for VLA and foundation-model research; (2) a free mobile app that lets any user record egocentric activity; and (3) a 200-hour dataset of diverse, long-form egocentric data with persistent state tracking across 584 sessions. We further show this data is a usable training signal:mid-training a VLA on it lowers held-out action-prediction error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。