arXiv:2607.13125cs.CVcs.AI2026-07被引 1

用小预算打造强理解的多模态生成模型,性能逼近闭源系统。

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

论文配图:Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
图 1 · 摘自论文原文
  • 通过强化多模态编码器和智能提示重写提升理解能力
  • 仅用2086万张图片训练,理论成本约40万美元
  • 适合关注开源多模态模型与低成本高效训练的研究者

我们提出Boogu-Image-0.1,一个开源的统一多模态理解与生成模型家族,包含Base、Turbo、Edit和Edit-Turbo四种变体。该模型在高质量文生图、快速推理、指令式编辑及中英文双语文本渲染方面表现优异。相较于依赖系统级集成的闭源模型(如Nano-Banana-Pro、GPT-Image-2),其内部机制不透明。本工作表明,通过增强多模态编码器、引入代理式提示重写等技术,并优化数据质量、训练流程与推理时扩展策略,即使在极低算力预算下也能显著提升生成与编辑性能。全面评估显示,Boogu-Image-0.1在标准基准上持续达到或超越其他开源模型,结果接近领先闭源系统。尤为关键的是,仅使用208.62百万张唯一图像,基础模型理论训练成本约为40万美元。我们公开分享实践经验、权重、代码与训练方案,采用Apache 2.0许可,以推动统一多模态理解与生成的开源生态发展。代码见:https://github.com/Boogu-Project/Boogu-Image。

原文摘要 · Abstract (English)

We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that strengthening the understanding capability of the system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.

多模态生成开源模型低成本训练图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。