让英文图像生成模型理解中文文化语义,生成更真实有文化的中文主题图像。
CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities
- 用多模态扩散变换器直接控制英文模型,提升中文语义理解。
- 在不重训模型前提下,中文生成质量显著优于现有方法。
- 兼容LoRA、ControlNet等插件,适合中文内容创作者使用。
我们提出中国文本适配器-流(CTA-Flux),一种将中文文本输入适配到原本基于英文语料训练的强大多媒体生成模型Flux的方法。尽管Flux在英文提示下具备出色图像生成能力,但面对非英文提示时表现不佳,主要由于训练数据以英语为主带来的语言与文化偏差。现有方法如翻译或微调双语映射,难以保留文化特异性语义,影响生成图像的真实性与质量。为此,我们提出新方法,将中文语义理解融入以英语为中心的文生图模型体系。相比依赖ControlNet类架构需大量参数且难以直接控制中文语义的方法,CTA-Flux采用多模态扩散变换器(MMDiT)直接控制Flux主干,大幅减少参数量,同时增强对中文语义的理解。该集成在无需大规模重训练的前提下,显著提升生成质量与文化真实性,且保持与现有文生图插件(如LoRA、IP-Adapter、ControlNet)的兼容性。实证评估表明,CTA-Flux可支持中英文提示,在图像生成质量、视觉真实性和中文语义忠实度方面均表现优异。
原文摘要 · Abstract (English)
We proposed the Chinese Text Adapter-Flux (CTA-Flux). An adaptation method fits the Chinese text inputs to Flux, a powerful text-to-image (TTI) generative model initially trained on the English corpus. Despite the notable image generation ability conditioned on English text inputs, Flux performs poorly when processing non-English prompts, particularly due to linguistic and cultural biases inherent in predominantly English-centric training datasets. Existing approaches, such as translating non-English prompts into English or finetuning models for bilingual mappings, inadequately address culturally specific semantics, compromising image authenticity and quality. To address this issue, we introduce a novel method to bridge Chinese semantic understanding with compatibility in English-centric TTI model communities. Existing approaches relying on ControlNet-like architectures typically require a massive parameter scale and lack direct control over Chinese semantics. In comparison, CTA-flux leverages MultiModal Diffusion Transformer (MMDiT) to control the Flux backbone directly, significantly reducing the number of parameters while enhancing the model's understanding of Chinese semantics. This integration significantly improves the generation quality and cultural authenticity without extensive retraining of the entire model, thus maintaining compatibility with existing text-to-image plugins such as LoRA, IP-Adapter, and ControlNet. Empirical evaluations demonstrate that CTA-flux supports Chinese and English prompts and achieves superior image generation quality, visual realism, and faithful depiction of Chinese semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。