arXiv:2501.17811cs.AIcs.CL2025-01被引 854

Janus-Pro通过数据与模型扩展,实现更强的多模态理解与图文生成能力。

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

  • 优化训练策略,扩大数据集,模型规模更大
  • 在图文指令遵循和生成稳定性上显著提升
  • 适合研究多模态生成与大模型应用的开发者

本文提出Janus-Pro,是前序工作Janus的升级版本。具体改进包括:(1) 优化的训练策略,(2) 扩展的训练数据,(3) 模型规模的扩大。借助这些改进,Janus-Pro在多模态理解与文本到图像指令遵循能力上取得显著进展,同时提升了文本到图像生成的稳定性。我们希望该工作能推动该领域的进一步探索。代码与模型已公开。

原文摘要 · Abstract (English)

In this work, we introduce Janus-Pro, an advanced version of the previous work Janus. Specifically, Janus-Pro incorporates (1) an optimized training strategy, (2) expanded training data, and (3) scaling to larger model size. With these improvements, Janus-Pro achieves significant advancements in both multimodal understanding and text-to-image instruction-following capabilities, while also enhancing the stability of text-to-image generation. We hope this work will inspire further exploration in the field. Code and models are publicly available.

多模态图文生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。