arXiv:2511.08168cs.AI2025-11

日本自研开源文生图模型oboro,用有限合规数据生成高质量图像。

oboro: Text-to-Image Synthesis on Limited Data using Flow-based Diffusion Transformer with MMH Attention

  • 基于流式扩散变换器与多模态混合注意力机制,适配小规模数据集
  • 仅用版权合规数据训练,仍可生成高保真图像
  • 首个日本自主开发的商用级开源文生图模型,代码权重公开

本项目是日本经产省(METI)与新能源产业技术综合开发机构(NEDO)支持的‘后5G信息通信系统基础设施强化研发计划——竞争性生成式AI基础模型(GENIAC)开发’的第二阶段课题。为应对日本动漫产业人力短缺问题,本项目从零开始研发图像生成模型。报告详细阐述了所开发的图像生成模型「oboro:」的技术规格。oboro: 是一款完全从零构建的图像生成模型,训练数据仅使用版权已清理的图像。其架构设计可在数据量有限的情况下生成高质量图像。基础模型权重与推理代码随报告一并公开。该项目标志着日本首次发布完全自主研发、面向商业应用的开源文生图AI。该模型源自开放源代码社区,通过保持开发过程透明,旨在推动日本人工智能研究者与工程师社区的发展,促进本土AI生态建设。

原文摘要 · Abstract (English)

This project was conducted as a 2nd-term adopted project of the "Post-5G Information and Communication System Infrastructure Enhancement R&D Project Development of Competitive Generative AI Foundation Models (GENIAC)," a business of the Ministry of Economy, Trade and Industry (METI) and the New Energy and Industrial Technology Development Organization (NEDO). To address challenges such as labor shortages in Japan's anime production industry, this project aims to develop an image generation model from scratch. This report details the technical specifications of the developed image generation model, "oboro:." We have developed "oboro:," a new image generation model built from scratch, using only copyright-cleared images for training. A key characteristic is its architecture, designed to generate high-quality images even from limited datasets. The foundation model weights and inference code are publicly available alongside this report. This project marks the first release of an open-source, commercially-oriented image generation AI fully developed in Japan. AiHUB originated from the OSS community; by maintaining transparency in our development process, we aim to contribute to Japan's AI researcher and engineer community and promote the domestic AI development ecosystem.

文生图日本自研小样本开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。