arXiv:2603.02803cs.CV2026-03被引 3

提升古希腊学术文本的版式感知识别能力,解决复杂注释结构识别难题。

Structure-Aware Text Recognition for Ancient Greek Critical Editions

  • 构建合成数据集与真实扫描基准,支持版式感知文本识别研究
  • 零样本下多数模型表现差于传统软件,但Qwen3VL-8B达1.0%中位字符错误率
  • 为历史学术文献数字化提供可扩展的视觉语言模型解决方案

视觉语言模型(VLMs)在端到端文档理解方面取得显著进展,但在处理历史学术文本复杂的版式语义方面仍存在局限。本文针对古希腊学术文本的结构感知识别问题,提出两种新资源:(i) 基于TEI/XML源生成的18.5万张页面图像的大型合成语料库,具有可控的排版与字体变化;(ii) 跨越一个世纪编辑与排版实践的真实扫描版本基准。利用这些数据集,我们在零样本与微调两种设置下评估了三种前沿VLMs。实验表明,当前VLM架构在高度结构化的历史文档面前存在明显不足;零样本条件下,多数模型性能显著低于现有成熟软件。然而,Qwen3VL-8B在真实扫描数据上达到1.0%的中位字符错误率,表现最佳。结果揭示了当前VLM在结构感知识别上的局限与未来潜力。

原文摘要 · Abstract (English)

Recent advances in visual language models (VLMs) have transformed end-to-end document understanding. However, their ability to interpret the complex layout semantics of historical scholarly texts remains limited. This paper investigates structure-aware text recognition for Ancient Greek critical editions, which have dense reference hierarchies and extensive marginal annotations. We introduce two novel resources: (i) a large-scale synthetic corpus of 185,000 page images generated from TEI/XML sources with controlled typographic and layout variation, and (ii) a curated benchmark of real scanned editions spanning more than a century of editorial and typographic practices. Using these datasets, we evaluate three state-of-the-art VLMs under both zero-shot and fine-tuning regimes. Our experiments reveal substantial limitations in current VLM architectures when confronted with highly structured historical documents. In zero-shot settings, most models significantly underperform compared to established off-the-shelf software. Nevertheless, the Qwen3VL-8B model achieves state-of-the-art performance, reaching a median Character Error Rate of 1.0\% on real scans. These results highlight both the current shortcomings and the future potential of VLMs for structure-aware recognition of complex scholarly documents.

文本识别古希腊文视觉语言模型版式理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。