仅用54万参数的极简策略,就能在LIBERO上达到95%成功率。
MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?

- 设计极简视觉-动作策略MINERVA,测试任务所需最小模型容量。
- 0.54M参数模型在2000次试错中平均成功率达95.1%,仅比大模型低2.4点。
- 模型只需少量参数即可高效运行,适合资源受限的机器人部署。
当前基于千亿参数的视觉-语言-动作(VLA)模型主导了LIBERO机械臂操作基准测试,但该任务实际所需的模型能力仍不明确。本文提出MINERVA(极小高效机器人视觉-动作策略),一系列刻意精简的视觉运动策略,用于测量此任务特定的容量下限。一个仅含0.54M参数的策略,在四个标准LIBERO任务套件上进行2000次滚动试验,平均成功率达到95.1%,仅比报告的LeRobot $/pi_{0.5}$ 结果低2.4分,而参数量少7,700倍。性能在接近100万参数时趋于饱和,低于0.25M则急剧下降。在广泛架构、训练和推理实验中,仅有动作块长度和视觉容量能持续超出±1分的种子波动范围。流匹配与直接L1回归相比无显著优势,且回归在GPU上最快可达3.8倍。任务ID置换探针显示,标准LIBERO指令条件主要依赖记忆任务:仅改变任务ID映射,成功率即降至随机水平。相同方法在89个LIBERO-90任务上达94.6%成功,而LIBERO-Plus扰动使性能降至46–56%,对光照变化几乎无鲁棒性。0.54M模型每控制步重规划耗时5–9毫秒(笔记本CPU),比SmolVLA快113倍,比$/\eta_{0.5}$快1,400倍,无需GPU。这些结果首次为LIBERO任务建立了实证容量下限,推动面向部署效率的容量感知设计与模型压缩。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark, but the model capacity actually required by the benchmark remains unclear. We introduce MINERVA (MINimal Efficient Robotic Vision-Action policy), a family of deliberately compact visuomotor policies designed to measure this task-specific capacity floor. A 0.54M-parameter policy achieves 95.1% average success over 2,000 rollouts on the four standard LIBERO suites, only 2.4 points below the reported LeRobot $\pi_{0.5}$ result despite using 7,700$\times$ fewer parameters. Performance saturates near 1M parameters and collapses below 0.25M. Across broad architectural, training, and inference sweeps, only action-chunk length and vision capacity consistently exceed a $\pm$1-point training-seed band. Flow matching provides no detectable advantage over direct L1 regression across three seeds, while regression is up to 3.8$\times$ faster on GPU. A task-ID permutation probe shows that standard LIBERO instruction conditioning primarily selects among memorized tasks: changing only the task-ID mapping reduces success to near chance. The same recipe achieves 94.6% success across 89 LIBERO-90 tasks, while LIBERO-Plus perturbations reduce performance to 46--56%, with near-zero robustness to photometric shifts. The 0.54M policy replans every control step in 5--9 ms per chunk on a laptop CPU, 113$\times$ faster than SmolVLA and 1,400$\times$ faster than $\pi_{0.5}$, without a GPU. These results establish a first empirical estimate of LIBERO's task-specific capacity floor and motivate capacity-aware design and distillation for deployment-efficient robot policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。