无需OCR和大量标注数据,实现多语言文本高保真合成
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
- 基于DiT架构,不依赖OCR编码器提取视觉特征
- 仅用1000样本以下即可在新语言上达到优秀性能
- 训练数据仅需竞品的1%,支持灵活多行布局控制
基于扩散模型的场景文本合成进展迅速,但现有方法通常依赖额外的视觉条件模块,并需要大规模标注数据支持多语言生成。本文重新审视复杂辅助模块的必要性,探索利用扩散模型固有的上下文推理能力,同时保证字形准确性和高保真场景融合。为此,提出TextFlux——一种基于DiT的多语言场景文本合成框架。其优势包括:(1) 无OCR模型架构,消除对专门提取文本特征的视觉条件模块需求;(2) 强大的多语言可扩展性,在低资源环境下表现优异,新增语言仅需少于1,000样本即能取得良好效果;(3) 简化训练设置,训练数据仅为竞争方法的1%;(4) 支持可控多行文本生成,具备精确的行级控制能力,优于仅支持单行或固定布局的方法。大量实验与可视化表明,TextFlux在定性与定量评估中均超越现有方法。
原文摘要 · Abstract (English)
Diffusion-based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large-scale annotated data to support multilingual generation. In this work, we revisit the necessity of complex auxiliary modules and further explore an approach that simultaneously ensures glyph accuracy and achieves high-fidelity scene integration, by leveraging diffusion models' inherent capabilities for contextual reasoning. To this end, we introduce TextFlux, a DiT-based framework that enables multilingual scene text synthesis. The advantages of TextFlux can be summarized as follows: (1) OCR-free model architecture. TextFlux eliminates the need for OCR encoders (additional visual conditioning modules) that are specifically used to extract visual text-related features. (2) Strong multilingual scalability. TextFlux is effective in low-resource multilingual settings, and achieves strong performance in newly added languages with fewer than 1,000 samples. (3) Streamlined training setup. TextFlux is trained with only 1% of the training data required by competing methods. (4) Controllable multi-line text generation. TextFlux offers flexible multi-line synthesis with precise line-level control, outperforming methods restricted to single-line or rigid layouts. Extensive experiments and visualizations demonstrate that TextFlux outperforms previous methods in both qualitative and quantitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。