arXiv:2503.06472cs.CVcs.MM2025-03ICCV

用视觉语言模型解决汉字书法的上下文理解难题

CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model

  • 分字切片+视觉文本对齐,精准提取与排序书法字符
  • 在全页识别中优于现有方法,接近专业人类水平
  • 适合研究书法智能分析或跨模态理解的学者

汉字书法作为联合国教科文组织遗产,因视觉模糊和文化复杂性而难以计算建模。现有AI系统无法有效上下文化其复杂笔迹,主要受限于标注数据不足和视觉-语义对齐差。本文提出CalliReader,一种视觉语言模型(VLM),通过三项创新解决汉字书法上下文化(CC²)问题:(1) 字符级切片实现精确字符提取与排序;(2) CalliAlign实现视觉-文本令牌压缩与对齐;(3) 嵌入式指令微调(e-IT)提升对齐效果并缓解数据稀缺。同时构建了首个面向整页书法上下文理解的基准CalliBench,解决了以往OCR与VQA方法在上下文碎片化、浅层推理和幻觉方面的三大缺陷。大量实验(含用户研究)验证CalliReader在页面级书法识别与解释上显著优于当前最优方法,甚至媲美专业人类,在准确率更高同时降低幻觉。与推理模型对比凸显准确识别是可靠理解的前提。定量分析表明模型高效,多文档与真实场景基准测试证实其强泛化能力。

原文摘要 · Abstract (English)

Chinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fail to contextualize their intricate scripts, because of limited annotated data and poor visual-semantic alignment. We propose CalliReader, a vision-language model (VLM) that solves the Chinese Calligraphy Contextualization (CC$^2$) problem through three innovations: (1) character-wise slicing for precise character extraction and sorting, (2) CalliAlign for visual-text token compression and alignment, (3) embedding instruction tuning (e-IT) for improving alignment and addressing data scarcity. We also build CalliBench, the first benchmark for full-page calligraphic contextualization, addressing three critical issues in previous OCR and VQA approaches: fragmented context, shallow reasoning, and hallucination. Extensive experiments including user studies have been conducted to verify our CalliReader's \textbf{superiority to other state-of-the-art methods and even human professionals in page-level calligraphy recognition and interpretation}, achieving higher accuracy while reducing hallucination. Comparisons with reasoning models highlight the importance of accurate recognition as a prerequisite for reliable comprehension. Quantitative analyses validate CalliReader's efficiency; evaluations on document and real-world benchmarks confirm its robust generalization ability.

书法识别视觉语言模型上下文理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。