arXiv:2502.10999cs.CVcs.AI2025-02EMNLP被引 8

无需字体标签,用扩散模型实现多语言文本的可控字体渲染

ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations

  • 用分割掩码在像素空间自监督学习字体特征
  • 零样本跨字体跨语言编辑成功实现
  • 适合需要灵活定制文本视觉样式的研究与开发

本工作证明,仅使用原始图像而无需字体标注,扩散模型即可实现可控制的多语言文本渲染。视觉文本渲染仍是重大挑战:尽管近期方法依赖字形条件扩散,但大规模真实数据集中无法获取精确字体标注,导致用户无法指定字体。为此,我们提出一种数据驱动方案,将条件扩散模型与文本分割模型结合,利用分割掩码在像素空间自监督捕捉并表示字体,完全消除对真实标签的需求,使用户可自由选择任意多语言字体进行文本渲染。实验验证了算法在零样本文本与字体编辑上的可行性,覆盖多种字体与语言,为实现通用视觉文本渲染提供了重要启示。代码已开源:github.com/bowen-upenn/ControlText。

原文摘要 · Abstract (English)

This work demonstrates that diffusion models can achieve font-controllable multilingual text rendering using just raw images without font label annotations.Visual text rendering remains a significant challenge. While recent methods condition diffusion on glyphs, it is impossible to retrieve exact font annotations from large-scale, real-world datasets, which prevents user-specified font control. To address this, we propose a data-driven solution that integrates the conditional diffusion model with a text segmentation model, utilizing segmentation masks to capture and represent fonts in pixel space in a self-supervised manner, thereby eliminating the need for any ground-truth labels and enabling users to customize text rendering with any multilingual font of their choice. The experiment provides a proof of concept of our algorithm in zero-shot text and font editing across diverse fonts and languages, providing valuable insights for the community and industry toward achieving generalized visual text rendering. Code is available at github.com/bowen-upenn/ControlText.

文本渲染扩散模型多语言可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。