用密集语言标注提升机器人策略学习效率,无需新数据采集。
How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning

- 通过多维度语言重标注,从已有视频中挖掘更多任务信号。
- 在RoboCasa上提升成功率5个百分点,接近专用标注的性能。
- 适合想低成本扩展机器人泛化能力的研究者与工程师。
机器人策略学习的规模化受示范数据收集成本制约,而现有示范的语言标注成本相对较低。本文研究语言密度作为提升固定机器人或自视角视频语料信号量的杠杆。提出DeMiAn(密集多维度标注)方法,分两阶段进行:首先利用视觉语言模型对示范片段生成物理运动、场景构成、机械臂姿态和推理四个互补维度的标注;再训练一个指导模型,在部署时将任务描述与初始场景快照映射为合适标注,异步运行以隐藏生成延迟。在超过100万条机器人操作视频和5万条EgoVerse人类自视角视频上验证,该方法在不新增示范数据的情况下,提升了视觉-语言-动作策略与基于视频的世界-动作模型性能。在RoboCasa上,指导模型使成功率比仅任务描述基线提升5点,距每任务专用标注的最优表现仅差3点。不同任务无主导性标注维度,表明选择合适的密集语言描述至关重要。此外,该方法还增强了组合任务与分布外泛化能力,并在中段训练和训练后阶段均推动了算力-性能边界,计入标注生成计算量后仍具优势。结果表明,密集重标注是机器人策略学习可落地的规模化路径。
原文摘要 · Abstract (English)
Scaling robot policy learning is bottlenecked by the cost of collecting demonstrations, while language annotations for existing demonstrations are comparatively cheap. We study language density as a lever for extracting more signal from a fixed robot or egocentric-video corpus. We introduce DeMiAn (Dense Multi-aspect Annotation), a two-stage approach that first re-labels demonstration segments with VLM-generated annotations along four complementary aspects: physical motion, scene composition, arm pose, and reasoning. A learned instructor then maps a task description and initial scene snapshot to a task-appropriate annotation at deployment, running asynchronously so generation latency is hidden behind policy execution. Across over 1M robot manipulation clips and 50K EgoVerse human-egocentric videos, DeMiAn improves both a vision-language-action policy and a video-based world-action model without collecting new demonstrations. On RoboCasa, the instructor raises success by 5 points over a task-only baseline and comes within 3 points of a per-task oracle. No fixed annotation aspect dominates across tasks, showing that selecting the right dense language matters. DeMiAn also improves composite-task and out-of-distribution performance, and shifts the compute-performance frontier in both mid-training and post-training after accounting for annotation-generation FLOPs. These results position dense re-annotation as a practical scaling lever for robot policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。