NeoBabel让多语言图像生成更准更快,打破英语垄断。
NeoBabel: A Multilingual Open Tower for Visual Generation
- 用大规模多语言预训练+高分辨率指令微调,直接提升跨语言生成能力。
- 在6种语言上表现超越现有模型,中文等非英语任务提升0.11以上。
- 模型小2-4倍,开源全套工具和数据集,助力公平生成AI研究。
文本到图像生成长期以英语为中心,加剧数字不平等。现有翻译管道导致语义偏差、计算开销与文化错位。我们提出NeoBabel,一个支持英、中、荷、法、印地、波斯六种语言的多语言图像生成框架,在性能、效率与包容性上达到新平衡。模型通过大规模多语言预训练与高分辨率指令微调训练。为评估能力,我们扩展两个英文基准为多语言版本:m-GenEval与m-DPG。NeoBabel在多语言任务上达0.75(m-GenEval)与0.68(m-DPG),优于基于多语言基座大模型的领先系统,且在英文任务上保持同等水平。其目标对齐训练有效增强跨语言泛化能力。我们还引入两项新指标,评估多语言对齐与混合编码提示鲁棒性。值得注意的是,尽管模型体积仅为2-4倍于同类系统,仍可匹敌甚至超越英文专用模型。我们开源完整工具包,含代码、模型检查点、1.24亿条多语言图文对数据集及标准化评估协议,推动包容性生成AI发展。结果表明,多语言能力并非权衡,而是提升鲁棒性、效率与文化准确性的催化剂。
原文摘要 · Abstract (English)
Text-to-image generation advancements have been predominantly English-centric, creating barriers for non-English speakers and perpetuating digital inequities. While existing systems rely on translation pipelines, these introduce semantic drift, computational overhead, and cultural misalignment. We introduce NeoBabel, a novel multilingual image generation framework that sets a new Pareto frontier in performance, efficiency and inclusivity, supporting six languages: English, Chinese, Dutch, French, Hindi, and Persian. The model is trained using a combination of large-scale multilingual pretraining and high-resolution instruction tuning. To evaluate its capabilities, we expand two English-only benchmarks to multilingual equivalents: m-GenEval and m-DPG. NeoBabel achieves state-of-the-art multilingual performance while retaining strong English capability, scoring 0.75 on m-GenEval and 0.68 on m-DPG. Notably, it performs on par with leading models on English tasks while outperforming them by +0.11 and +0.09 on multilingual benchmarks, even though these models are built on multilingual base LLMs. This demonstrates the effectiveness of our targeted alignment training for preserving and extending crosslingual generalization. We further introduce two new metrics to rigorously assess multilingual alignment and robustness to code-mixed prompts. Notably, NeoBabel matches or exceeds English-only models while being 2-4x smaller. We release an open toolkit, including all code, model checkpoints, a curated dataset of 124M multilingual text-image pairs, and standardized multilingual evaluation protocols, to advance inclusive AI research. Our work demonstrates that multilingual capability is not a trade-off but a catalyst for improved robustness, efficiency, and cultural fidelity in generative AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。