arXiv:2605.15298cs.ROcs.AI2026-05

从人类第一视角视频中提取物理常识,提升机器人跨域泛化能力

PhysBrain 1.0 Technical Report

论文配图:PhysBrain 1.0 Technical Report
图 1 · 摘自论文原文
  • 将人类视频转为带问答的物理常识监督信号
  • 在多模态与具身控制任务上达顶尖性能,跨域表现尤佳
  • 适合关注机器人物理理解与迁移学习的研究者

视觉-语言-动作模型发展迅速,但仅靠机器人轨迹难以全面学习物理知识。PhysBrain 1.0提出新路径:将大规模人类第一视角视频转化为结构化的物理常识监督信号,再用于机器人适配。数据引擎提取场景元素、空间动态、动作执行及深度感知关系,生成可用于训练视觉-语言模型(VLM)的问答监督数据。所获物理先验通过保持能力且敏感于语言的适配设计,迁移至视觉-语言-动作策略中。在多模态问答与具身控制基准测试中,包括ERQA、PhysBench、SimplerEnv-WidowX、LIBERO和RoboCasa,PhysBrain 1.0均取得当前最优(SOTA)结果,尤其在SimplerEnv上表现出显著的跨域泛化能力。结果表明,从人类交互视频中扩展物理常识,可有效连接多模态理解与机器人行动。

原文摘要 · Abstract (English)

Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a complementary route: converting large-scale human egocentric video into structured physical commonsense supervision before robot adaptation. Our data engine extracts scene elements, spatial dynamics, action execution, and depth-aware relations, then turns them into question-answer supervision for training PhysBrain VLMs. The resulting physical priors are further transferred to VLA policies through a capability-preserving and language-sensitive adaptation design. Across multimodal QA benchmarks and embodied control benchmarks, including ERQA, PhysBench, SimplerEnv-WidowX, LIBERO, and RoboCasa, PhysBrain 1.0 achieves SOTA results and shows especially strong out-of-domain performance on SimplerEnv. These results suggest that scaling physical commonsense from human interaction video can provide an effective bridge from multimodal understanding to robot action.

物理常识视觉语言动作机器人学习跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。