arXiv:2603.15620cs.CVcs.RO2026-03中稿 · ECCV被引 15

提出动态操作新数据集与模型,显著提升机器人在移动目标环境中的泛化能力。

Towards Generalizable Robotic Manipulation in Dynamic Environments

  • 构建包含110万条轨迹的动态操作数据集DOMINO,支持多层级任务评估。
  • 提出的PUMA模型在动态任务上成功率提升6.3%,优于现有方法。
  • 动态数据训练可增强时空表征,还能迁移到静态任务中使用。

视觉-语言-动作(VLA)模型在静态操作中表现优异,但在存在移动目标的动态环境中性能下降明显。这一差距主要源于动态操作数据集稀缺,以及主流VLA依赖单帧观测,限制了其时空推理能力。为此,我们提出DOMINO——一个大规模动态操作数据集与基准,包含35个具有层次复杂性的任务、超过110,000条专家轨迹,以及多维度评估体系。通过系统实验,我们评估现有VLA在动态任务上的表现,探索有效提升动态感知的训练策略,并验证动态数据的泛化能力。此外,我们提出PUMA,一种动态感知的VLA架构:通过引入以场景为中心的历史光流和专用世界查询,隐式预测对象中心的未来状态,实现历史感知与短时预测的耦合。结果表明,PUMA在动态任务中取得当前最优性能,成功率达6.3%绝对提升。更重要的是,动态数据训练可生成鲁棒的时空表示,并能迁移至静态任务。所有代码与数据已公开于https://github.com/H-EmbodVis/DOMINO。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap primarily stems from a scarcity of dynamic manipulation datasets and the reliance of mainstream VLAs on single-frame observations, restricting their spatiotemporal reasoning capabilities. To address this, we introduce DOMINO, a large-scale dataset and benchmark for generalizable dynamic manipulation, featuring 35 tasks with hierarchical complexities, over 110K expert trajectories, and a multi-dimensional evaluation suite. Through comprehensive experiments, we systematically evaluate existing VLAs on dynamic tasks, explore effective training strategies for dynamic awareness, and validate the generalizability of dynamic data. Furthermore, we propose PUMA, a dynamics-aware VLA architecture. By integrating scene-centric historical optical flow and specialized world queries to implicitly forecast object-centric future states, PUMA couples history-aware perception with short-horizon prediction. Results demonstrate that PUMA achieves state-of-the-art performance, yielding a 6.3% absolute improvement in success rate over baselines. Moreover, we show that training on dynamic data fosters robust spatiotemporal representations that transfer to static tasks. All code and data are available at https://github.com/H-EmbodVis/DOMINO.

机器人操作动态感知视觉语言模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。