首个端到端生成整行手写的扩散模型,兼顾字形与风格一致性。
DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
- 用双正则化变分自编码器分离字形与书写风格
- 在真实手写数据集上达到最高字形准确率与风格保真度
- 适合需要高效生成连贯手写文本的场景
深度生成模型已推动文本到在线手写生成(TOHG)的发展,旨在根据文本输入和风格参考合成逼真的笔迹轨迹。然而,现有方法多聚焦于字符或单词级生成,导致全行文本生成效率低且缺乏整体结构建模。为此,我们提出 DiffInk,首个用于整行手写生成的潜在扩散变换器框架。首先引入 InkVAE,一种增强的序列变分自编码器,包含两项互补的潜在空间正则化损失:(1) 基于OCR的损失,确保字形级别准确性;(2) 风格分类损失,保留书写风格。该双重正则化使潜在空间具备语义结构,实现字符内容与写作风格的有效解耦。随后提出 InkDiT,一种新型潜在扩散变换器,整合目标文本与参考风格以生成连贯笔迹轨迹。实验表明,DiffInk 在字形准确率与风格保真度上均超越现有最先进方法,同时显著提升生成效率。
原文摘要 · Abstract (English)
Deep generative models have advanced text-to-online handwriting generation (TOHG), which aims to synthesize realistic pen trajectories conditioned on textual input and style references. However, most existing methods still primarily focus on character- or word-level generation, resulting in inefficiency and a lack of holistic structural modeling when applied to full text lines. To address these issues, we propose DiffInk, the first latent diffusion Transformer framework for full-line handwriting generation. We first introduce InkVAE, a novel sequential variational autoencoder enhanced with two complementary latent-space regularization losses: (1) an OCR-based loss enforcing glyph-level accuracy, and (2) a style-classification loss preserving writing style. This dual regularization yields a semantically structured latent space where character content and writer styles are effectively disentangled. We then introduce InkDiT, a novel latent diffusion Transformer that integrates target text and reference styles to generate coherent pen trajectories. Experimental results demonstrate that DiffInk outperforms existing state-of-the-art (SOTA) methods in both glyph accuracy and style fidelity, while significantly improving generation efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。