构建首个覆盖97455个汉字的超大规模识别数据集,推动中文字符识别研究
MegaHan97K: A Large-Scale Dataset for Mega-Category Chinese Character Recognition with over 97K Categories
- 构建覆盖最新国标标准的97455类汉字数据集,规模为现有最大数据集的6倍
- 通过手写、历史和合成三类子集平衡长尾分布,提升模型泛化能力
- 揭示超大类别场景下的存储、形似字识别与零样本学习新挑战,适合相关研究者使用
中文字符是中华文化的基础,其类别数量极其庞大且持续扩展,最新国家标准GB18030-2022包含87,887个字符。准确识别如此海量的字符(即超类别识别)对文化遗产保护与数字应用至关重要。尽管光学字符识别(OCR)技术进展显著,但因缺乏全面数据集,超类别识别仍处于空白状态,现有最大数据集仅含16,151个类别。为此,我们提出MegaHan97K,一个覆盖前所未有的97,455个汉字类别的超大规模数据集。主要贡献包括:(1) 首个完整支持最新GB18030-2022标准的数据集,类别数至少是现有数据集的六倍;(2) 通过手写、历史和合成三个子集有效缓解长尾分布问题,实现各类别样本均衡;(3) 全面基准测试揭示了超类别场景中的新挑战,如存储压力增大、形似字识别困难、零样本学习难题,同时也为未来研究开辟巨大空间。据我们所知,MegaHan97K可能是当前OCR领域乃至更广泛模式识别领域中类别数量最多的数据集。数据集已开源:https://github.com/SCUT-DLVCLab/MegaHan97K。
原文摘要 · Abstract (English)
Foundational to the Chinese language and culture, Chinese characters encompass extraordinarily extensive and ever-expanding categories, with the latest Chinese GB18030-2022 standard containing 87,887 categories. The accurate recognition of this vast number of characters, termed mega-category recognition, presents a formidable yet crucial challenge for cultural heritage preservation and digital applications. Despite significant advances in Optical Character Recognition (OCR), mega-category recognition remains unexplored due to the absence of comprehensive datasets, with the largest existing dataset containing merely 16,151 categories. To bridge this critical gap, we introduce MegaHan97K, a mega-category, large-scale dataset covering an unprecedented 97,455 categories of Chinese characters. Our work offers three major contributions: (1) MegaHan97K is the first dataset to fully support the latest GB18030-2022 standard, providing at least six times more categories than existing datasets; (2) It effectively addresses the long-tail distribution problem by providing balanced samples across all categories through its three distinct subsets: handwritten, historical and synthetic subsets; (3) Comprehensive benchmarking experiments reveal new challenges in mega-category scenarios, including increased storage demands, morphologically similar character recognition, and zero-shot learning difficulties, while also unlocking substantial opportunities for future research. To the best of our knowledge, the MetaHan97K is likely the dataset with the largest classes not only in the field of OCR but may also in the broader domain of pattern recognition. The dataset is available at https://github.com/SCUT-DLVCLab/MegaHan97K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。