用第一视角视频训练机器人,实现语言指令下的灵巧操作。
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

- 从真实场景第一视角视频构建9.6小时高质量数据集,提升训练效率9倍。
- 模型在40多个任务中成功执行自由指令,长时任务成功率超75%。
- 适合研究灵巧操作、具身智能与视觉-语言-动作对齐的开发者使用。
可操控性是通用机器人策略的核心能力,但灵巧手系统因缺乏大规模、语言对齐且动作精准的示范数据而难以实现。为此,我们提出一个全栈系统,通过第一人称人类视频规模化预训练灵巧手视觉-语言-动作模型,并支持高效实机微调。系统集成EgoSmith数据流水线,将野外第一视角视频转化为9.6千小时高质量预训练数据,吞吐量比现有最佳方法高9倍,准确率更优;统一机器人框架支持远程操控与人机协同修正;以及基于世界模型增强的EgoSteer VLA,在优化基础设施上训练完成。人类数据预训练赋予EgoSteer语言引导的操作先验,经机器人微调与DAgger优化后进一步提升。实验证明,EgoSteer能稳健执行40多个多样化任务中的自由指令,具备故障恢复、灵巧操作与泛化能力。预训练模型还可少样本适配复杂长时任务(如折纸),在两种机器人本体上成功率超75%。系统、数据与模型已开源:https://egosteer.github.io/。
原文摘要 · Abstract (English)
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training. It integrates EgoSmith, a data pipeline that curates in-the-wild egocentric videos into 9.6K hours of high-quality pre-training data with 9x higher throughput and better accuracy than prior SOTA; a unified robot stack for teleoperation and human-in-the-loop correction; and EgoSteer, a world-model-enhanced VLA trained on optimized infrastructure. Human-data pre-training equips EgoSteer with language-guided manipulation priors, which are grounded through robot post-training and improved by DAgger refinement. Empirically, EgoSteer robustly executes free-form instructions across 40+ diverse tasks, demonstrating failure recovery, dexterity, and generalization. The pre-trained model also few-shot adapts to complex long-horizon tasks, including box folding, on two embodiments with 75+% success. We open-source the system, data, and model at https://egosteer.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。