开源70亿到720亿参数的具身智能脑模型,性能媲美顶级闭源系统。
Pelican-VL 1.0: A Foundation Brain Model for Embodied Intelligence
- 用元循环学习框架训练,通过刻意练习实现智能进化
- 在1000+ A800 GPU上训练,每检查点超5万小时算力,性能比基础模型提升20.3%
- 适合研究具身智能、多模态机器人和自主学习系统的开发者
本报告介绍Pelican-VL 1.0,一个参数规模从70亿到720亿的开源具身智能脑模型家族。其使命是将强大智能嵌入各类实体中。该模型目前为最大规模的开源多模态具身智能模型。核心优势在于数据力量与智能自适应学习机制的深度融合。通过metaloop从包含40亿以上标记的原始数据集中提炼出高质量数据集。模型在超过1000个A800 GPU的大规模集群上训练,每个检查点消耗超过5万小时的A800 GPU算力。相比基础模型,性能提升20.3%,且在知名具身任务基准上超越1000亿级开源模型10.6%,达到领先闭源系统的水平。我们提出一种新框架DPPO(深思熟虑的实践策略优化),受人类元认知启发,具体表现为一个‘强化学习-精炼-诊断-监督微调’的元循环流程。
原文摘要 · Abstract (English)
This report presents Pelican-VL 1.0, a new family of open-source embodied brain models with parameter scales ranging from 7 billion to 72 billion. Our explicit mission is clearly stated as: To embed powerful intelligence into various embodiments. Pelican-VL 1.0 is currently the largest-scale open-source embodied multimodal brain model. Its core advantage lies in the in-depth integration of data power and intelligent adaptive learning mechanisms. Specifically, metaloop distilled a high-quality dataset from a raw dataset containing 4+ billion tokens. Pelican-VL 1.0 is trained on a large-scale cluster of 1000+ A800 GPUs, consuming over 50k+ A800 GPU-hours per checkpoint. This translates to a 20.3% performance uplift from its base model and outperforms 100B-level open-source counterparts by 10.6%, placing it on par with leading proprietary systems on well-known embodied benchmarks. We establish a novel framework, DPPO (Deliberate Practice Policy Optimization), inspired by human metacognition to train Pelican-VL 1.0. We operationalize this as a metaloop that teaches the AI to practice deliberately, which is a RL-Refine-Diagnose-SFT loop.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。