用扩散模型从手写图像恢复笔迹轨迹,提升精度与泛化能力。
Handwriting Trajectory Recovery with Diffusion Models

- 基于扩散模型建模笔迹生成,以图像条件驱动轨迹恢复。
- 在CASIA-OLHWDB上实现更优的时间相似性与形状保真度。
- 可跨文字类型迁移,训练中文能恢复拉丁字母笔顺。
从离线手写图像中恢复在线笔迹轨迹(即笔画恢复),是离线转在线转换任务,适用于笔画级编辑和司法鉴定。本文首次提出基于扩散模型的框架解决该问题。方法将轨迹恢复建模为图像条件生成,使用去噪扩散模型采样与墨迹一致的笔迹轨迹。在CASIA-OLHWDB(1.0–1.1)上的大量定量评估表明,该方法即使对复杂多笔画字符也能实现高精度恢复,显著优于PEN-Net、Cross-VAE等代表性方法,在时间相似性(DTW/LDTW)和形状保真度(AIoU)上均有提升。此外,模型能捕捉普遍的书写顺序规律,具备良好泛化能力,例如:仅在中文字符上训练的模型可部分合理恢复拉丁字母的书写顺序。
原文摘要 · Abstract (English)
Recovering online pen trajectories from offline handwriting images, often referred to as handwriting trajectory recovery (stroke recovery), is an offline-to-online conversion task with applications in stroke-level editing and forensic analysis. We propose, to the best of our knowledge, the first diffusion-model-based framework for this task. Our method formulates trajectory recovery as image-conditioned generation and uses a denoising diffusion model to sample pen trajectories consistent with the observed ink trace. Through extensive quantitative evaluations on CASIA-OLHWDB (1.0-1.1), we verify that the proposed approach enables accurate recovery even for complex multi-stroke characters, substantially improving both temporal similarity (DTW/LDTW) and shape fidelity (AIoU) over representative prior methods such as PEN-Net and Cross-VAE. We further show that the model captures general stroke-order tendencies and generalizes to classes unseen during training, exemplified by cross-script transfer: a model trained on Chinese characters can recover reasonable stroke orders for Latin letters to some extent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。