为楚简文字设计多模态多粒度分词器,提升古文字识别准确率
Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slip Scripts
- 先检测字符边界,再分级识别字符与部件
- 在10万级标注数据集上,词性标注F1提升5.5%
- 适合古文字研究者与跨模态语言模型开发者
本研究针对春秋战国时期(公元前771-256年)的楚简文字(CBS)设计了一种多模态多粒度分词器。鉴于古汉字复杂的层级结构——单个字符可能由多个部件组成——该分词器首先通过字符检测定位边界,再在字符与子字符两个层面进行识别。为支持学术研究,我们构建了首个大规模CBS标注数据集,包含超过10万张字符图像。基于该数据集的词性标注任务中,使用本分词器相较主流子词分词器实现5.5%的相对F1分数提升。本工作不仅有助于深化对特定古文字的研究,也为其他古代汉语文字研究提供了潜在技术支撑。
原文摘要 · Abstract (English)
This study presents a multi-modal multi-granularity tokenizer specifically designed for analyzing ancient Chinese scripts, focusing on the Chu bamboo slip (CBS) script used during the Spring and Autumn and Warring States period (771-256 BCE) in Ancient China. Considering the complex hierarchical structure of ancient Chinese scripts, where a single character may be a combination of multiple sub-characters, our tokenizer first adopts character detection to locate character boundaries, and then conducts character recognition at both the character and sub-character levels. Moreover, to support the academic community, we have also assembled the first large-scale dataset of CBSs with over 100K annotated character image scans. On the part-of-speech tagging task built on our dataset, using our tokenizer gives a 5.5% relative improvement in F1-score compared to mainstream sub-word tokenizers. Our work not only aids in further investigations of the specific script but also has the potential to advance research on other forms of ancient Chinese scripts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。