arXiv:2608.15812cs.CV2026-08

用真人笔迹匹配代替生成,实现个性化中文手写逼真还原

From Generation to Matching: A Development Report on Personalized Chinese Handwriting

论文配图:From Generation to Matching: A Development Report on Personalized Chinese Handwriting
图 1 · 摘自论文原文
  • 从生成转为匹配:基于真实笔迹库逐字寻找最相似人类样本
  • 197个汉字全部找到对应真人笔迹,100字测试集100%命中
  • 适合需要高保真手写模拟的场景,如数字签名、个性化字体

本文记录了一项关于个性化中文手写的冻结工程实践。项目始于约200张单用户真实手写图像,覆盖197个独特汉字,最初以少样本生成未见字形为目标。但多轮尝试发现:增强结构规范性会削弱个性特征,过度个性化则破坏关键笔画识别。因此转向真实人类笔迹等价类重构任务,利用多作者CASIA笔迹库发现多数用户笔迹已有真实对应样本。任务由生成转为字符级匹配,并通过跨作者组合生成虚拟写作者。系统采用真实墨迹特征、字符专属人群百分位、前20候选剪枝及贪心最难优先整行选择策略。在目标字符集上,197个用户汉字均有真实人类候选;100字评估子集实现100/100覆盖。已知字留出测试中,某行被视觉判断几乎无法与真实用户笔迹区分。60轮稳定性审计显示所有结果均落在预设的A级机器代理区域,但非独立人类评分。最终证据支持稳定实用的B级质量,部分输出在用户定义标准下接近A级。

原文摘要 · Abstract (English)

This paper documents a frozen engineering project on personalized Chinese handwriting. The project started from approximately 200 real handwriting images from one user, covering 197 unique Chinese characters, and was initially formulated as few-shot generation of unseen characters. A sequence of canonical-centered personalization routes repeatedly exposed the same conflict: increasing structural pressure made outputs more canonical, while increasing personalization could damage identity-defining strokes. The project was therefore reset around real-human character equivalence classes. A multi-writer CASIA candidate pool showed that a USER-compatible realization often already existed among valid human samples. The task consequently changed from synthesis to character-wise matching, followed by cross-writer composition into a virtual writer. The frozen system uses real-ink features, character-specific human population percentiles, top-20 candidate pruning, and greedy hardest-first whole-row selection. On the covered target set, all 197 USER characters had real-human candidates, and the 100-character evaluation subset was covered 100/100. Knowncharacter held-out comparisons included a row judged visually almost indistinguishable from genuine USER handwriting. A 60- episode stability audit placed every episode in a predefined A-like machine-proxy region, but these were not independent human A-level judgments. The final evidence supports stable practical B-level quality, with many outputs approaching A-level under the USER-defined criterion. The report records why generation became unnecessary for this case without claiming unrestricted or universal handwriting synthesis.

手写生成个性化笔迹匹配中文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。