统一生成与识别中文书法,兼顾字形结构与整体布局。
UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy
- 联合训练生成与识别任务,互为约束提升质量
- 在8000+样本上实现最优的连笔与排版一致性
- 适用于甲骨文等古文字,适合文化遗产数字化
中文书法的计算复现仍具挑战。现有方法或仅生成高质量单字而忽略页面级美感(如连笔、间距),或试图合成整页却牺牲书法规范性。我们提出 extbf{UniCalli},一个面向列级生成与识别的统一扩散框架。通过联合训练:识别任务约束生成器保持字形结构,生成任务提供风格与布局先验,二者协同形成概念级抽象,显著提升有限数据下的性能。我们构建了超8000份数字化作品数据集,其中约4000份密集标注。UniCalli采用非对称加噪与栅格化框图作为空间先验,基于合成、标注及未标注数据混合训练。模型在生成质量、连笔连续性与布局保真度方面达到当前最优,同时识别能力更强。框架成功扩展至甲骨文与埃及象形文字。代码与数据见https://github.com/EnVision-Research/UniCalli。
原文摘要 · Abstract (English)
Computational replication of Chinese calligraphy remains challenging. Existing methods falter, either creating high-quality isolated characters while ignoring page-level aesthetics like ligatures and spacing, or attempting page synthesis at the expense of calligraphic correctness. We introduce \textbf{UniCalli}, a unified diffusion framework for column-level recognition and generation. Training both tasks jointly is deliberate: recognition constrains the generator to preserve character structure, while generation provides style and layout priors. This synergy fosters concept-level abstractions that improve both tasks, especially in limited-data regimes. We curated a dataset of over 8,000 digitized pieces, with ~4,000 densely annotated. UniCalli employs asymmetric noising and a rasterized box map for spatial priors, trained on a mix of synthetic, labeled, and unlabeled data. The model achieves state-of-the-art generative quality with superior ligature continuity and layout fidelity, alongside stronger recognition. The framework successfully extends to other ancient scripts, including Oracle bone inscriptions and Egyptian hieroglyphs. Code and data can be viewed in \href{https://github.com/EnVision-Research/UniCalli}{this URL}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。