arXiv:2505.18479cs.CV2025-05CVPR

用三维法向量增强合成数据,让文字更自然地融入真实场景。

Syn3DTxt: Embedding 3D Cues for Scene Text Generation

  • 在2D数据中加入表面法向量,引入三维几何信息。
  • 新数据集显著提升复杂3D场景下的文字渲染效果。
  • 适合做场景文字生成、三维内容合成的研究者参考。

本研究针对合成数据集中三维上下文不足的问题展开探索。尽管扩散模型等技术提升了场景文字生成的某些方面,但多数方法仍依赖2D数据,训练样本主要来自电影海报和书籍封面,难以捕捉真实场景中空间布局与视觉效果的复杂交互。传统2D数据缺乏必要的几何线索,无法准确实现文字在多样背景中的嵌入。为此,我们提出一种新的合成数据构建标准,通过在常规2D数据中加入表面法向量,丰富三维场景特征。该方法旨在增强空间关系表征,为未来场景文字渲染提供更稳健的基础。大量实验表明,依据此标准构建的数据集能显著改善几何上下文表达,推动复杂3D空间条件下文字渲染的进一步发展。

原文摘要 · Abstract (English)

This study aims to investigate the challenge of insufficient three-dimensional context in synthetic datasets for scene text rendering. Although recent advances in diffusion models and related techniques have improved certain aspects of scene text generation, most existing approaches continue to rely on 2D data, sourcing authentic training examples from movie posters and book covers, which limits their ability to capture the complex interactions among spatial layout and visual effects in real-world scenes. In particular, traditional 2D datasets do not provide the necessary geometric cues for accurately embedding text into diverse backgrounds. To address this limitation, we propose a novel standard for constructing synthetic datasets that incorporates surface normals to enrich three-dimensional scene characteristic. By adding surface normals to conventional 2D data, our approach aims to enhance the representation of spatial relationships and provide a more robust foundation for future scene text rendering methods. Extensive experiments demonstrate that datasets built under this new standard offer improved geometric context, facilitating further advancements in text rendering under complex 3D-spatial conditions.

文本生成三维建模合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。