arXiv:2605.14708cs.CV2026-05中稿 · CVPR被引 4

跨语言场景文字生成,精准保持风格一致性。

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation

论文配图:StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
图 1 · 摘自论文原文
  • 双分支编码器提取多语言文字风格特征。
  • 新损失函数提升字符间风格连贯性。
  • 掩码引导推理实现生成与参考文字对齐。

风格条件化场景文字生成在复杂背景中面临精确提取文本风格并保持字符间细粒度风格一致性的挑战,尤其在多语言脚本中更为突出。本文提出StyleTextGen框架,通过学习感知和复现不同语言与书写系统的视觉文字风格。方法包括:首先设计双分支风格编码器,实现复杂真实场景下鲁棒的多语言文字风格表征;其次引入文本风格一致性损失,增强风格连贯性并提升整体视觉质量;第三构建掩码引导推理策略,确保生成文字与参考文字间的精准风格对齐。为系统评估,我们构建了包含单语言与跨语言设置的双语场景文字风格基准StyleText-CE。大量实验表明,StyleTextGen在风格一致性和跨语言泛化能力上显著优于现有方法,建立了多语言风格条件化文字生成的新SOTA性能。

原文摘要 · Abstract (English)

Style-conditioned scene text generation faces unique challenges in extracting precise text styles from complex backgrounds and maintaining fine-grained style consistency across characters, especially for multilingual scripts. We propose StyleTextGen, a novel framework that learns to perceive and replicate visual text styles across different languages and writing systems. Our approach features three key contributions: First, we introduce a dual-branch style encoder dedicated to style modeling, yielding robust multilingual text style representations in complex real-world scenes. Second, we design a text style consistency loss that enhances style coherence and improves overall visual quality. Third, we develop a mask-guided inference strategy that ensures precise style alignment between generated and reference text. To facilitate systematic evaluation, we construct StyleText-CE, a bilingual scene text style benchmark covering both monolingual and cross-lingual settings. Extensive experiments demonstrate that StyleTextGen significantly outperforms existing methods in style consistency and cross-lingual generalization, establishing new state-of-the-art performance in multilingual style-conditioned text generation.

文字生成跨语言风格控制视觉生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。