arXiv:2501.05329cs.LGcs.RO2025-01中稿 · AAMAS 2025被引 1

将大模型知识压缩到小模型,实现高效多任务强化学习部署

Knowledge Transfer in Model-Based Reinforcement Learning Agents for Efficient Multi-Task Learning

  • 用知识蒸馏将317M参数大模型压缩为1M参数小模型
  • 在MT30基准上得分28.45,远超原小模型的18.93
  • 支持低资源环境部署,适合机器人等实际场景

我们提出一种高效的模型基于强化学习中的知识迁移方法,解决在资源受限环境下部署大型世界模型的挑战。该方法将高容量多任务智能体(317M参数)蒸馏为仅1M参数的紧凑模型,在MT30基准上取得了28.45的归一化分数,显著优于原始1M参数模型的18.93分,证明了该蒸馏技术有效整合复杂多任务知识。此外,我们采用FP16后训练量化,使模型大小减少50%的同时保持性能。本工作弥合了大模型能力与实际部署限制之间的差距,为机器人及其他资源受限领域提供了可扩展的高效多任务强化学习解决方案。

原文摘要 · Abstract (English)

We propose an efficient knowledge transfer approach for model-based reinforcement learning, addressing the challenge of deploying large world models in resource-constrained environments. Our method distills a high-capacity multi-task agent (317M parameters) into a compact 1M parameter model, achieving state-of-the-art performance on the MT30 benchmark with a normalized score of 28.45, a substantial improvement over the original 1M parameter model's score of 18.93. This demonstrates the ability of our distillation technique to consolidate complex multi-task knowledge effectively. Additionally, we apply FP16 post-training quantization, reducing the model size by 50% while maintaining performance. Our work bridges the gap between the power of large models and practical deployment constraints, offering a scalable solution for efficient and accessible multi-task reinforcement learning in robotics and other resource-limited domains.

强化学习知识蒸馏多任务学习模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。