构建统一甲骨文字表,助力古汉语模型训练与数字化
InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling
- 整合未编码甲骨文与古今汉字,建立统一字符列表
- 在甲骨文语料上训练模型,任务性能显著提升
- 适合古汉语NLP、考古语言学研究者使用
构建历史语言模型对辅助考古溯源和理解古代文化至关重要。然而现有资源在训练有效模型方面面临两大挑战:一是历史语料稀缺,基于大规模文本的无监督学习效率低下;二是古代文字存在时间跨度大、演变复杂,缺乏完整的字符编码体系,阻碍了古文字的数字化与计算处理,尤其影响早期中文书写。为此,我们提出InteChar——一个统一且可扩展的字符列表,融合未编码的甲骨文与传统及现代汉字,实现历史文本的一致数字化与表示,为古文字建模奠定基础。为评估其有效性,我们构建了以甲骨文为核心、结合专家标注与大模型增强的数据集OracleCS。大量实验表明,在OracleCS上使用InteChar训练的模型在多项历史语言理解任务中均取得显著提升,验证了该方法的有效性,并为未来古汉语自然语言处理研究提供坚实基础。
原文摘要 · Abstract (English)
Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effective LMs on historical texts. First, the scarcity of historical language samples renders unsupervised learning approaches based on large text corpora highly inefficient, hindering effective pre-training. Moreover, due to the considerable temporal gap and complex evolution of ancient scripts, the absence of comprehensive character encoding schemes limits the digitization and computational processing of ancient texts, particularly in early Chinese writing. To address these challenges, we introduce InteChar, a unified and extensible character list that integrates unencoded oracle bone characters with traditional and modern Chinese. InteChar enables consistent digitization and representation of historical texts, providing a foundation for robust modeling of ancient scripts. To evaluate the effectiveness of InteChar, we construct the Oracle Corpus Set (OracleCS), an ancient Chinese corpus that combines expert-annotated samples with LLM-assisted data augmentation, centered on Chinese oracle bone inscriptions. Extensive experiments show that models trained with InteChar on OracleCS achieve substantial improvements across various historical language understanding tasks, confirming the effectiveness of our approach and establishing a solid foundation for future research in ancient Chinese NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。