让机器人理解任务进展,自动判断何时该停止或调整动作。
ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation
- 用预训练模型从视频和文本中估算任务进展,准确率达0.07误差。
- 通过可微分的进展引导机制,提升长序列操作的成功率。
- 适合做复杂、多步骤的机器人操控任务,尤其在真实场景中表现强。
现有视觉-语言-动作(VLA)模型在机器人操作中缺乏任务进展感知,通常依赖人工设计的终止规则,这在包含多个子目标的长程任务中尤为严重。本文提出新模型ProgressVLA,技术贡献有二:(1) 鲁棒的进展估计:在大规模无监督视频-文本机器人数据集上预训练进展估计算法,在仿真中预测残差低至0.07(范围[0,1]),并实现对未见真实样本的零样本泛化;(2) 可微分的进展引导:引入逆动力学世界模型,将动作标记映射为未来隐空间状态,再经进展估计算法处理,通过最大进展正则化建立可微管道,实现进展驱动的动作优化。在CALVIN与LIBERO基准上的大量实验及真实机器人部署均显示,其成功率和泛化能力显著优于强基线。
原文摘要 · Abstract (English)
Most existing vision-language-action (VLA) models for robotic manipulation lack progress awareness, typically relying on hand-crafted heuristics for task termination. This limitation is particularly severe in long-horizon tasks involving cascaded sub-goals. In this work, we investigate the estimation and integration of task progress, proposing a novel model named {\textbf \vla}. Our technical contributions are twofold: (1) \emph{robust progress estimation}: We pre-train a progress estimator on large-scale, unsupervised video-text robotic datasets. This estimator achieves a low prediction residual (0.07 on a scale of $[0, 1]$) in simulation and demonstrates zero-shot generalization to unseen real-world samples, and (2) \emph{differentiable progress guidance}: We introduce an inverse dynamics world model that maps predicted action tokens into future latent visual states. These latents are then processed by the progress estimator; by applying a maximal progress regularization, we establish a differentiable pipeline that provides progress-piloted guidance to refine action tokens. Extensive experiments on the CALVIN and LIBERO benchmarks, alongside real-world robot deployment, consistently demonstrate substantial improvements in success rates and generalization over strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。