构建首个多属性汉字书法数据集,支持风格、朝代、书家等细粒度识别。
MCCD: A Multi-Attribute Chinese Calligraphy Character Dataset Annotated with Script Styles, Dynasties, and Calligraphers
- 构建7765类共32.9万张独立书法图像,标注字体、朝代、书家三类属性
- 在10种字体、15个朝代、142位书家的子集上建立基准性能
- 适用于书法识别、作者辨识与汉字演变研究,适合文化数字化领域
书法属性研究(如字体、朝代、书家)具有重要文化与历史价值。然而,汉字书法风格随朝代更迭与书家个性演变显著,导致字符及其属性识别难度极高。现有书法数据集稀少,且大多仅提供字符级标注,缺乏附加属性信息,严重制约了深入研究。为此,我们提出首个多属性汉字书法字符数据集(MCCD),包含7,765类别、总计329,715张孤立书法图像,并基于字体(10类)、朝代(15时期)、书家(142人)三类属性构建三个子集。丰富的多属性标注使其适用于书法字符识别、作者辨识及汉字演变研究。我们在MCCD及其所有子集上开展单任务与多任务识别实验,结果表明笔画结构复杂性及属性间交互显著提升了识别难度。MCCD填补了高质量书法数据集的空白,为推动书法研究及跨领域应用提供关键资源。数据集开源:https://github.com/SCUT-DLVCLab/MCCD。
原文摘要 · Abstract (English)
Research on the attribute information of calligraphy, such as styles, dynasties, and calligraphers, holds significant cultural and historical value. However, the styles of Chinese calligraphy characters have evolved dramatically through different dynasties and the unique touches of calligraphers, making it highly challenging to accurately recognize these different characters and their attributes. Furthermore, existing calligraphic datasets are extremely scarce, and most provide only character-level annotations without additional attribute information. This limitation has significantly hindered the in-depth study of Chinese calligraphy. To fill this gap, we present a novel Multi-Attribute Chinese Calligraphy Character Dataset (MCCD). The dataset encompasses 7,765 categories with a total of 329,715 isolated image samples of Chinese calligraphy characters, and three additional subsets were extracted based on the attribute labeling of the three types of script styles (10 types), dynasties (15 periods) and calligraphers (142 individuals). The rich multi-attribute annotations render MCCD well-suited diverse research tasks, including calligraphic character recognition, writer identification, and evolutionary studies of Chinese characters. We establish benchmark performance through single-task and multi-task recognition experiments across MCCD and all of its subsets. The experimental results demonstrate that the complexity of the stroke structure of the calligraphic characters, and the interplay between their different attributes, leading to a substantial increase in the difficulty of accurate recognition. MCCD not only fills a void in the availability of detailed calligraphy datasets but also provides valuable resources for advancing research in Chinese calligraphy and fostering advancements in multiple fields. The dataset is available at https://github.com/SCUT-DLVCLab/MCCD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。