让文字在复杂曲面上精准贴合方向,提升电商广告图像生成质量
OrienText: Surface Oriented Textual Image Generation
- 用表面法向量作为条件输入,引导文本在曲面上正确朝向
- 自建数据集上显著优于现有方法,文字贴合度与方向准确性更高
- 适合需要精确图文融合的电商、广告和影视场景
图像中的文本内容在电商领域至关重要,尤其在营销活动、产品展示、广告及娱乐产业中。当前基于扩散模型的文本到图像生成方法虽能产出高质量图像,但在处理建筑构件、横幅或墙壁等具有不同视角的复杂曲面时,常难以准确放置文本。本文提出面向表面方向的文本图像生成方法(OrienText),利用区域特定的表面法向量作为扩散模型的条件输入,确保文本在图像上下文中的精准渲染与正确朝向。我们在自构建的数据集上验证了该方法的有效性,并与现有文本图像生成方法进行对比。
原文摘要 · Abstract (English)
Textual content in images is crucial in e-commerce sectors, particularly in marketing campaigns, product imaging, advertising, and the entertainment industry. Current text-to-image (T2I) generation diffusion models, though proficient at producing high-quality images, often struggle to incorporate text accurately onto complex surfaces with varied perspectives, such as angled views of architectural elements like buildings, banners, or walls. In this paper, we introduce the Surface Oriented Textual Image Generation (OrienText) method, which leverages region-specific surface normals as conditional input to T2I generation diffusion model. Our approach ensures accurate rendering and correct orientation of the text within the image context. We demonstrate the effectiveness of the OrienText method on a self-curated dataset of images and compare it against the existing textual image generation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。