arXiv:2603.27670cs.ROcs.AI2026-03被引 10

让机器人理解任务进展,自动判断何时该停止或调整动作。

ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation

  • 用预训练模型从视频和文本中估算任务进展,准确率达0.07误差。
  • 通过可微分的进展引导机制,提升长序列操作的成功率。
  • 适合做复杂、多步骤的机器人操控任务,尤其在真实场景中表现强。

现有视觉-语言-动作(VLA)模型在机器人操作中缺乏任务进展感知,通常依赖人工设计的终止规则,这在包含多个子目标的长程任务中尤为严重。本文提出新模型ProgressVLA,技术贡献有二:(1) 鲁棒的进展估计:在大规模无监督视频-文本机器人数据集上预训练进展估计算法,在仿真中预测残差低至0.07(范围[0,1]),并实现对未见真实样本的零样本泛化;(2) 可微分的进展引导:引入逆动力学世界模型,将动作标记映射为未来隐空间状态,再经进展估计算法处理,通过最大进展正则化建立可微管道,实现进展驱动的动作优化。在CALVIN与LIBERO基准上的大量实验及真实机器人部署均显示,其成功率和泛化能力显著优于强基线。

原文摘要 · Abstract (English)

Most existing vision-language-action (VLA) models for robotic manipulation lack progress awareness, typically relying on hand-crafted heuristics for task termination. This limitation is particularly severe in long-horizon tasks involving cascaded sub-goals. In this work, we investigate the estimation and integration of task progress, proposing a novel model named {\textbf \vla}. Our technical contributions are twofold: (1) \emph{robust progress estimation}: We pre-train a progress estimator on large-scale, unsupervised video-text robotic datasets. This estimator achieves a low prediction residual (0.07 on a scale of $[0, 1]$) in simulation and demonstrates zero-shot generalization to unseen real-world samples, and (2) \emph{differentiable progress guidance}: We introduce an inverse dynamics world model that maps predicted action tokens into future latent visual states. These latents are then processed by the progress estimator; by applying a maximal progress regularization, we establish a differentiable pipeline that provides progress-piloted guidance to refine action tokens. Extensive experiments on the CALVIN and LIBERO benchmarks, alongside real-world robot deployment, consistently demonstrate substantial improvements in success rates and generalization over strong baselines.

机器人操作进展感知扩散策略视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。