用单字生成高精度书法,让笔触更真实流畅。
InkDiffuser: High-Fidelity One-shot Chinese Calligraphy via Differentiable Morphological Optimization

- 通过高频特征融合与可微墨迹损失,精准还原笔画轮廓。
- 仅需一个参考字,生成字体在结构和细节上均优于现有方法。
- 适合对书法生成真实感要求高的艺术创作与数字人文研究。
当前中文书法生成方法存在笔画渲染不佳、墨迹形态失真等问题,导致输出视觉质量与艺术流畅性不足。为此,我们提出基于扩散模型的单样本书法生成框架InkDiffuser。为保障高保真渲染,引入两项核心创新:高频增强机制与可微墨迹结构(DIS)损失。受个体样本中高频信息通常携带轮廓细节的启发,通过显式融合高频表征提升内容提取精度,实现更准确的字体结构建模。此外,提出一种将可微形态操作融入扩散过程的DIS损失,使模型能学习墨迹结构的显式分解,从而精细化优化笔画轮廓,显著提升生成书法的视觉真实感。在多种书体及复杂汉字上的大量实验表明,InkDiffuser仅需单个参考字即可生成高质量书法字体,在结构一致性、细节保真度与视觉真实性方面均优于现有少样本字体生成方法。代码已开源:https://github.com/JingVIPLab/InkDiffuser。
原文摘要 · Abstract (English)
Current Chinese calligraphy generation methods suffer from poor stroke rendering and unrealistic ink morphology, resulting in outputs with limited visual fidelity and artistic fluidity. To address this problem, we propose \textbf{InkDiffuser}, a diffusion-based generative framework for one-shot Chinese calligraphy synthesis. To guarantee high-fidelity rendering, we introduce two core contributions: a high-frequency enhancement mechanism and a Differentiable Ink Structure (DIS) loss that explicitly regularizes ink morphology. Inspired by the observation that high-frequency information in individual samples typically carries contour details, we enhance content extraction by explicitly fusing high-frequency representations for more accurate font structure. Furthermore, we propose a differentiable ink structure loss that integrates differentiable morphological operations into the diffusion process. By allowing the model to learn an explicit decomposition of ink-trace structures, DIS facilitates fine-grained refinement of stroke contours and delivers significantly improved visual realism in the generated calligraphy. Extensive experiments on various calligraphic styles and complex characters demonstrate that InkDiffuser can generate superior calligraphy fonts with realistic ink rendering effects from only a single reference glyph and outperform existing few-shot font generation approaches in structural consistency, detail fidelity, and visual authenticity. The code is available at the following address: https://github.com/JingVIPLab/InkDiffuser.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。