arXiv:2608.26673cs.RO2026-08

用0.68万参数实现高效机器人操作,靠预测感知运动模型而非直接映射。

PredVLA: Predictive Sensorimotor Modeling for Sub-Million-Parameter Robot Manipulation

论文配图:PredVLA: Predictive Sensorimotor Modeling for Sub-Million-Parameter Robot Manipulation
图 1 · 摘自论文原文
  • 通过分层递归预测视觉和本体感觉,减少对数据预训练依赖。
  • 在LIBERO上达86.9%成功率,比同类模型高3.7到7.4倍。
  • 预测路径缺失导致性能下降约70%,证明其核心优势。

大型预训练视觉-语言-动作模型在机器人操作中表现优异,而紧凑型模型通常通过压缩观测到动作的范式来提升效率。我们探究预测性感知运动建模是否能在有限参数下比直接映射更有效。本文提出PredVLA,一个仅含0.68百万可训练参数、无需机器人数据预训练的语言条件预测编码策略。其分层递归动态预测视觉特征与本体感知,观测仅通过预测误差驱动的在线推断影响隐状态。在LIBERO基准上,PredVLA在三个短时程任务套件中取得86.9%的平均成功率,全部四个套件为75.4%。在使用相同冻结前端、示范数据、动作解码器与评估协议的控制对比中,相比参数匹配的Transformer与LSTM行为克隆策略,成功率分别提升3.7倍与7.4倍。机制级过渡分析表明,将预测路径替换为直接观测输入导致最大性能下降,约占端点差距的70%。进一步消融实验揭示了训练时隐状态推断、测试时误差回归、分层时间尺度及感官预测误差通道的独立贡献。这些结果支持预测性感知运动建模作为小型语言条件机器人控制的强大归纳偏置。

原文摘要 · Abstract (English)

Large pretrained vision-language-action models achieve strong robot-manipulation performance, while compact alternatives have largely pursued efficiency by compressing the prevailing observation-to-action paradigm. We investigate whether predictive sensorimotor modeling can make more effective use of a limited parameter budget than direct observation-to-action mapping. We present PredVLA, a language-conditioned predictive-coding policy with only 0.68 million trainable network parameters and no robot-data pretraining. Its hierarchical recurrent dynamics predict visual features and proprioception, while observations influence latent state only through prediction-error-driven online inference. On LIBERO, PredVLA achieves an 86.9% mean success rate across the three short-horizon suites and 75.4% across all four suites. Under a controlled comparison using the same frozen front end, demonstrations, action decoder, and evaluation protocol, PredVLA achieves 3.7x and 7.4x the three-suite mean success rates of parameter-matched Transformer and LSTM behavior-cloning policies, respectively. A mechanism-by-mechanism transition to the recurrent behavior-cloning baseline shows that replacing the predictive pathway with direct observation input produces the largest single performance drop, accounting for approximately $70\%$ of the endpoint gap. Further ablations identify distinct contributions from training-time latent inference, test-time error regression, hierarchical timescales, and sensory prediction-error channels. Together, these results support predictive sensorimotor modeling as a strong inductive bias for compact language-conditioned robot control.

机器人控制预测建模小模型语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。