无需训练即可实现图像中多语言文字的精准变形与融合。
DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images
- 分层编辑:分离文字与背景,实现可控几何变换。
- 复杂变换下仍保持视觉质量,大尺度旋转缩放效果优。
- 全免训练设计,适配多场景快速部署。
我们提出DanceText,一种无需训练的多语言图像文本编辑框架,支持复杂几何变换并实现前景与背景无缝融合。尽管基于扩散模型的生成方法在文本引导图像合成中表现良好,但在旋转、平移、缩放、扭曲等非平凡操作下常缺乏可控性且布局不一致。DanceText引入分层编辑策略,将文字与背景分离,实现模块化、可控制的几何变换;同时设计深度感知模块,对齐变换后文字与重建背景的外观与视角,提升真实感与空间一致性。该框架完全采用预训练模块,无需任务微调即可灵活部署。在AnyWord-3M基准上的大量实验表明,本方法在视觉质量上表现优异,尤其在大规模复杂变换场景中优势明显。代码已开源。
原文摘要 · Abstract (English)
We present DanceText, a training-free framework for multilingual text editing in images, designed to support complex geometric transformations and achieve seamless foreground-background integration. While diffusion-based generative models have shown promise in text-guided image synthesis, they often lack controllability and fail to preserve layout consistency under non-trivial manipulations such as rotation, translation, scaling, and warping. To address these limitations, DanceText introduces a layered editing strategy that separates text from the background, allowing geometric transformations to be performed in a modular and controllable manner. A depth-aware module is further proposed to align appearance and perspective between the transformed text and the reconstructed background, enhancing photorealism and spatial consistency. Importantly, DanceText adopts a fully training-free design by integrating pretrained modules, allowing flexible deployment without task-specific fine-tuning. Extensive experiments on the AnyWord-3M benchmark demonstrate that our method achieves superior performance in visual quality, especially under large-scale and complex transformation scenarios. Code is avaible at https://github.com/YuZhenyuLindy/DanceText.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。