液态神经网络+混合密度头,更小更快更稳的模仿学习新方案
Liquid Networks with Mixture Density Heads for Efficient Imitation Learning
- 用液态神经网络搭配混合密度头建模策略,实现轻量化多模态输出
- 参数减半(430万对860万),离线预测误差降2.4倍,推理快1.8倍
- 低数据和中等数据下表现更稳健,适合资源受限场景
我们在共享骨干网络的对比框架下,将液态神经网络结合混合密度头的方法与扩散策略在Push-T、RoboMimic Can和PointMaze任务上进行比较,严格控制输入、训练预算和评估设置,仅考察策略头的差异。实验显示,液态策略模型参数量约为430万(对比860万),离线预测误差降低2.4倍,推理速度提升1.8倍。在样本效率测试中(训练数据占比1%至46.42%),液态模型在低数据和中等数据条件下均表现出显著优势。闭环测试结果虽存在噪声,但整体趋势与离线排名一致,表明优秀的密度建模有助于部署表现,但不完全决定闭环成功。总体而言,液态递归多模态策略为模仿学习提供了一种紧凑且实用的替代方案。
原文摘要 · Abstract (English)
We compare liquid neural networks with mixture density heads against diffusion policies on Push-T, RoboMimic Can, and PointMaze under a shared-backbone comparison protocol that isolates policy-head effects under matched inputs, training budgets, and evaluation settings. Across tasks, liquid policies use roughly half the parameters (4.3M vs. 8.6M), achieve 2.4x lower offline prediction error, and run 1.8 faster at inference. In sample-efficiency experiments spanning 1% to 46.42% of training data, liquid models remain consistently more robust, with especially large gains in low-data and medium-data regimes. Closed-loop results on Push-T and PointMaze are directionally consistent with offline rankings but noisier, indicating that strong offline density modeling helps deployment while not fully determining closed-loop success. Overall, liquid recurrent multimodal policies provide a compact and practical alternative to iterative denoising for imitation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。