用人类第一视角视频训练机器人物理智能,提升规划与控制能力。
PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
- 将人类第一视角视频转为带证据和时序一致性的机器人动作监督数据
- 在E2E-3M数据集上训练出的PhysBrain显著提升机器人规划能力
- 适合研究视觉语言动作系统与机器人样本效率优化的学者
机器人泛化依赖物理智能:在第一人称感知与行动下,对状态变化、接触密集交互及长程规划的推理能力。视觉语言模型(VLM)是视觉-语言-动作(VLA)系统的核心,但依赖第三人称训练数据造成视角差距。收集大量机器人第一视角数据虽理想却因成本与多样性受限而不可行。相比之下,人类第一视角视频具有高可扩展性与丰富交互上下文,但身体结构差异使其无法直接使用。为此,我们提出一个第一视角到具身化转换管道,将原始人类第一视角视频转化为多层级、基于模板的具身化监督数据,强制证据锚定与时间一致性,实现了大规模埃戈中心到具身化数据集(E2E-3M)的构建。基于E2E-3M训练的具身化大脑PhysBrain,在第一视角理解方面表现显著提升,尤其在规划任务中。它提供了一种第一视角感知初始化,使VLA微调更高效,成功率更高,验证了从人类第一视角监督向下游机器人控制的有效迁移。
原文摘要 · Abstract (English)
Robotic generalization relies on physical intelligence: the ability to reason about state changes, contact-rich interactions, and long-horizon planning under egocentric perception and action. Vision Language Models (VLMs) are essential to Vision-Language-Action (VLA) systems, but the reliance on third-person training data creates a viewpoint gap for humanoid robots. Collecting massive robot-centric data is an ideal but impractical solution due to cost and diversity constraints. Conversely, human egocentric videos offer a highly scalable data source with rich interaction context, yet the embodiment mismatch prevents the direct application. To bridge this gap, we propose an Egocentric2Embodiment Translation Pipeline that transforms raw human egocentric videos into multi-level, schema-driven embodiment supervision with enforced evidence grounding and temporal consistency, enabling the construction of the Egocentric2Embodiment dataset (E2E-3M) at scale. An egocentric-aware embodied brain, termed PhysBrain, is obtained by training on the E2E-3M dataset. PhysBrain exhibits substantially improved egocentric understanding, particularly for planning. It provides an egocentric-aware initialization that enables more sample-efficient VLA fine-tuning and higher success rates, demonstrating effective transfer from human egocentric supervision to downstream robot control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。