arXiv:2603.07543cs.CVcs.MM2026-03中稿 · as oral presentati…被引 1

用分块对比与风格感知量化,实现单张手写样例的高质量生成。

CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization

  • 将手写风格建模为离散视觉令牌,分离出倾斜、笔画粗细等核心特征。
  • 在多个语言数据集上生成图像细节更丰富,真实感更强,优于现有方法。
  • 适合需要快速适配新书写风格的研究与应用,如个性化手写字体生成。

单样本风格化手写图像生成虽近年取得显著进展,但仍因仅凭一张参考图难以捕捉人类手写的复杂多样性而面临挑战。现有方法仍难以生成视觉上吸引人且逼真的手写图像,也难适应未见过的书写风格,无法有效分离不变风格特征(如倾斜度、笔画宽度、弧度)并忽略无关噪声。为此,我们提出基于去噪扩散模型的常量方法(CONSTANT),通过三大创新:1)风格感知量化(SAQ)模块,将风格建模为捕捉特定概念的离散视觉令牌;2)对比学习目标,确保嵌入空间中的令牌分布清晰且语义明确;3)潜在分块对比(LatentPCE)目标,通过对齐生成与真实特征在潜在空间中多尺度局部块,提升图像质量和局部结构。在包含英文、中文及我们提出的越南语维HTGen数据集在内的多个基准数据集上的大量实验表明,本方法在适配新参考风格和生成高细节图像方面均优于当前最优方法。代码已开源于GitHub。

原文摘要 · Abstract (English)

One-shot styled handwriting image generation, despite achieving impressive results in recent years, remains challenging due to the difficulty in capturing the intricate and diverse characteristics of human handwriting by using solely a single reference image. Existing methods still struggle to generate visually appealing and realistic handwritten images and adapt to complex, unseen writer styles, struggling to isolate invariant style features (e.g., slant, stroke width, curvature) while ignoring irrelevant noise. To tackle this problem, we introduce Patch Contrastive Enhancement and Style-Aware Quantization via Denoising Diffusion (CONSTANT), a novel one-shot handwriting generation via diffusion model. CONSTANT leverages three key innovations: 1) a Style-Aware Quantization (SAQ) module that models style as discrete visual tokens capturing distinct concepts; 2) a contrastive objective to ensure these tokens are well-separated and meaningful in the embedding style space; 3) a latent patch-based contrastive (LLatentPCE) objective help improving quality and local structures by aligning multiscale spatial patches of generated and real features in latent space. Extensive experiments and analysis on benchmark datasets from multiple languages, including English, Chinese, and our proposed ViHTGen dataset for Vietnamese, demonstrate the superiority of adapting to new reference styles and producing highly detailed images of our method over state-of-the-art approaches. Code is available at GitHub

手写生成扩散模型风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。