让机器人在有限算力下精准感知任务进度,提升复杂操作成功率。
ProgVLA: Progress-Aware Robot Manipulation Skill Learning

- 用双阶段视觉语言动作编码器压缩多模态序列,保留跨模态对齐。
- 0.1B参数模型在长任务中成功率超更大预训练模型。
- 适合资源受限场景下的复杂机械臂操作,尤其擅长多物体长序列任务。
我们提出ProgVLA,一种轻量级视觉-语言-动作(VLA)模型,专为在严格计算与内存限制下实现可靠机器人操作而设计。该模型通过显式建模长期任务中的进展,高效处理多模态长序列。其核心包含两部分:第一,采用两阶段的Perceiver重采样架构,将可变长度的视觉、语言和本体感觉流压缩为固定数量的控制就绪上下文标记,显著缩短序列长度并保持跨模态对齐;第二,通过离线强化学习目标训练一组辅助进度头,联合学习归一化剩余时域目标上的价值函数。这使策略具备内部任务进展估计能力,并支持优势加权与成功加权的流匹配模仿学习。在两个主流多任务机器人操作基准上,0.1B参数的ProgVLA模型在长时序和高难度任务层级上超越了显著更大的预训练基线模型。消融实验表明,所学上下文重采样模块和任务自适应视觉微调是最大贡献项,而进度感知训练在长时序和多物体任务上带来稳定增益。此外,我们在真实世界玩具厨房环境中验证了该方法的有效性。
原文摘要 · Abstract (English)
We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The model specifically focuses on efficiently processing long multi-modal sequences by maintaining an explicit representation of task progress over extended horizons. To this end, ProgVLA integrates two key components. First, a multi-modal encoder with a two-stage Perceiver resampling scheme compresses variable-length visual, language, and proprioceptive streams into a fixed set of control-ready context tokens, substantially reducing sequence length while preserving cross-modal grounding. Second, an auxiliary set of progress heads is trained with offline reinforcement learning (RL) objectives to jointly learn critics over normalized remaining-horizon targets. This provides the policy with an internal estimate of task progress and enables advantage- and success-weighted flow-matching imitation learning. On two well-established multi-task robot manipulation benchmarks, a 0.1B-parameter ProgVLA model reaches success rates that are competitive with, and on long-horizon and harder task tiers exceed, substantially larger pretrained baselines. Ablations indicate that the learned context resampler and task-adaptive visual fine-tuning are the largest single contributors, while progress-aware training provides a consistent additional gain that is concentrated on long-horizon and multi-object tasks. We further validate the approach in real-world toy-kitchen environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。