arXiv:2608.11741cs.CVcs.AI2026-08中稿 · the Dataset Track …

构建首个专家审核的古汉字释义数据集,推动多模态模型理解古文字演变。

JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis

论文配图:JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis
图 1 · 摘自论文原文
  • 设计四层递进的古汉字释义任务,涵盖形、义、源流分析
  • 创建超50万条问答对的数据集,通过专家模板和人工校验确保准确
  • 适合研究古文字、跨模态理解及历史语言学的学者与开发者

古汉字释义需融合视觉观察、语言分析与历史背景,但现有计算方法仅聚焦字符识别等子任务,缺乏系统性数据集与评估基准。为此,我们提出古汉字释义(ACCE)任务,包含四层渐进式分析:基础识别、字形分析、意义解释与历时演变。构建了两个互补资源:JieZi-Dataset是首个大规模、专家审核的视觉-语言问答训练数据集,含50万余条问答对,通过专家设计模板与文本引用控制生成,并在关键阶段进行人工验证以保证学术准确性;JieZi-Bench是与释义流程一致的评估基准,由人类专家构建并验证,四层均有权威辞书参考答案,且独立于训练数据。多模态大模型实验表明,当前模型在基础识别上表现良好,但在字形分析、语义推理和历时理解上仍困难。在JieZi-Dataset上微调后,各层级性能显著提升。代码与数据集已开源。

原文摘要 · Abstract (English)

The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis. To address this limitation, we introduce Ancient Chinese Character Exegesis (ACCE), a vision-language question answering (VQA) task that models the scholarly exegesis process. ACCE is organized into four progressive levels: basic character identification, glyph-form analysis, meaning exegesis, and diachronic evolution analysis. To support this task, we construct two complementary resources. JieZi-Dataset is the first large-scale, expert-audited VQA training dataset for ACCE, comprising over 500K QA pairs. It is constructed via a pipeline that reduces factual errors by constraining generation with expert-designed templates and source-text references. Human verification is further applied at each key stage to ensure scholarly accuracy. JieZi-Bench is an evaluation benchmark aligned with the exegesis process, constructed and verified by human experts to ensure evaluation reliability. It consists of four levels with reference answers curated from authoritative lexicographic works held separate from the training data. Experiments on multimodal large language models show that current models perform well on basic identification but struggle with glyph analysis, semantic reasoning, and diachronic understanding. Fine-tuning on JieZi-Dataset substantially improves performance across all four levels. Code and dataset are available at https://github.com/Ran00w/JieZi.

古汉字视觉问答多模态语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。