用无任务依赖动作数据训练机器人通用操作能力,提升跨任务泛化性。
AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation
- 通过自动采集无任务关联的动作数据,分离机械体态动态与高层策略
- 在多个操作任务中成功率提升30%-40%,测试准确率提高51%
- 适合需要跨平台、跨任务泛化的机器人视觉运动控制场景
机器人操作政策的可泛化性依赖于数据,但现有机器人操作数据稀缺且常与具体硬件绑定,导致跨任务和跨平台迁移困难。本文提出任务无关的体态建模方法,直接从任务无关的动作数据中学习体态动态,并将其与高层策略学习解耦。任务无关数据以独立图像-动作对形式存在,能覆盖整个体态工作空间,不同于依赖具体任务的序列化数据。该数据驱动方法规避了传统动力学建模局限,实现动作数据在不同任务间的可扩展复用。基于此,我们提出AnyPos统一框架,结合大规模自动化任务无关探索与基于逆动力学的学习的鲁棒体态建模。AnyPos能规模化生成多样且安全的轨迹,通过解耦手臂与末端执行器运动,并采用方向感知解码器,在分布偏移下稳定预测,可无缝对接多种高层策略模型。相比标准基线,AnyPos在测试准确率上提升51%;在操作微波炉、烤面包、叠衣服、浇花、刷盘子等任务中,成功率较强基线提升30%-40%。结果表明,数据驱动的体态建模是缓解数据稀缺、实现视觉运动控制跨任务与跨平台泛化的有效路径。
原文摘要 · Abstract (English)
Learning generalizable manipulation policies hinges on data, yet robot manipulation data is scarce and often entangled with specific embodiments, making both cross-task and cross-platform transfer difficult. We tackle this challenge with task-agnostic embodiment modeling, which learns embodiment dynamics directly from task-agnostic action data and decouples them from high-level policy learning. By focusing on exploring all feasible actions of the embodiment to capture what is physically feasible and consistent, task-agnostic data takes the form of independent image-action pairs with the potential to cover the entire embodiment workspace, unlike task-specific data, which is sequential and tied to concrete tasks. This data-driven perspective bypasses the limitations of traditional dynamics-based modeling and enables scalable reuse of action data across different tasks. Building on this principle, we introduce AnyPos, a unified pipeline that integrates large-scale automated task-agnostic exploration with robust embodiment modeling through inverse dynamics learning. AnyPos generates diverse yet safe trajectories at scale, then learns embodiment representations by decoupling arm and end-effector motions and employing a direction-aware decoder to stabilize predictions under distribution shift, which can be seamlessly coupled with diverse high-level policy models. In comparison to the standard baseline, AnyPos achieves a 51% improvement in test accuracy. On manipulation tasks such as operating a microwave, toasting bread, folding clothes, watering plants, and scrubbing plates, AnyPos raises success rates by 30-40% over strong baselines. These results highlight data-driven embodiment modeling as a practical route to overcoming data scarcity and achieving generalization across tasks and platforms in visuomotor control. Project page: https://embodiedfoundation.github.io/vidar_anypos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。