arXiv:2505.14683cs.CV2025-05被引 828

开源模型BAGEL实现图文视频统一理解与生成,具备复杂多模态推理能力。

Emerging Properties in Unified Multimodal Pretraining

论文配图:Emerging Properties in Unified Multimodal Pretraining
图 1 · 摘自论文原文
  • 基于万亿级跨模态数据训练的统一解码器模型
  • 在生成与理解任务上超越现有开源模型,支持自由图像编辑等能力
  • 适合研究多模态大模型与开放生态构建的开发者与学者

统一多模态理解和生成已在前沿闭源系统中展现出惊人能力。本文提出BAGEL,一个原生支持多模态理解和生成的开源基础模型。BAGEL是基于万亿级令牌的统一解码器模型,预训练数据来自大规模交错的文本、图像、视频和网页数据。在多样化多模态交错数据的规模扩展下,BAGEL展现出复杂的多模态推理能力。结果表明,其在标准基准测试中显著优于现有开源统一模型,在多模态生成与理解方面表现突出,具备自由图像操作、未来帧预测、3D操作和世界导航等高级能力。为促进多模态研究发展,我们公开关键发现、预训练细节、数据创建协议,并发布代码与检查点。项目页面见https://bagel-ai.org/

原文摘要 · Abstract (English)

Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. In this work, we introduce BAGEL, an open-source foundational model that natively supports multimodal understanding and generation. BAGEL is a unified, decoder-only model pretrained on trillions of tokens curated from large-scale interleaved text, image, video, and web data. When scaled with such diverse multimodal interleaved data, BAGEL exhibits emerging capabilities in complex multimodal reasoning. As a result, it significantly outperforms open-source unified models in both multimodal generation and understanding across standard benchmarks, while exhibiting advanced multimodal reasoning abilities such as free-form image manipulation, future frame prediction, 3D manipulation, and world navigation. In the hope of facilitating further opportunities for multimodal research, we share the key findings, pretraining details, data creation protocal, and release our code and checkpoints to the community. The project page is at https://bagel-ai.org/

多模态统一模型开源推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。