arXiv:2509.23624cs.CV2025-09中稿 · ICLR被引 3

首个端到端生成整行手写的扩散模型,兼顾字形与风格一致性。

DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation

  • 用双正则化变分自编码器分离字形与书写风格
  • 在真实手写数据集上达到最高字形准确率与风格保真度
  • 适合需要高效生成连贯手写文本的场景

深度生成模型已推动文本到在线手写生成(TOHG)的发展,旨在根据文本输入和风格参考合成逼真的笔迹轨迹。然而,现有方法多聚焦于字符或单词级生成,导致全行文本生成效率低且缺乏整体结构建模。为此,我们提出 DiffInk,首个用于整行手写生成的潜在扩散变换器框架。首先引入 InkVAE,一种增强的序列变分自编码器,包含两项互补的潜在空间正则化损失:(1) 基于OCR的损失,确保字形级别准确性;(2) 风格分类损失,保留书写风格。该双重正则化使潜在空间具备语义结构,实现字符内容与写作风格的有效解耦。随后提出 InkDiT,一种新型潜在扩散变换器,整合目标文本与参考风格以生成连贯笔迹轨迹。实验表明,DiffInk 在字形准确率与风格保真度上均超越现有最先进方法,同时显著提升生成效率。

原文摘要 · Abstract (English)

Deep generative models have advanced text-to-online handwriting generation (TOHG), which aims to synthesize realistic pen trajectories conditioned on textual input and style references. However, most existing methods still primarily focus on character- or word-level generation, resulting in inefficiency and a lack of holistic structural modeling when applied to full text lines. To address these issues, we propose DiffInk, the first latent diffusion Transformer framework for full-line handwriting generation. We first introduce InkVAE, a novel sequential variational autoencoder enhanced with two complementary latent-space regularization losses: (1) an OCR-based loss enforcing glyph-level accuracy, and (2) a style-classification loss preserving writing style. This dual regularization yields a semantically structured latent space where character content and writer styles are effectively disentangled. We then introduce InkDiT, a novel latent diffusion Transformer that integrates target text and reference styles to generate coherent pen trajectories. Experimental results demonstrate that DiffInk outperforms existing state-of-the-art (SOTA) methods in both glyph accuracy and style fidelity, while significantly improving generation efficiency.

手写生成扩散模型风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。