用专家路由解决满文低资源OCR中多书写风格识别难题
Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study

- 将迭代微调的检查点作为领域专家,通过页面级分类器按风格路由
- 在三个测试集上达到0.30%~4.83%的字符错误率,媲美专用专家
- 仅运行体专家是专为该领域训练,适合历史文本多风格低资源场景
满文历史文档需识别多种视觉差异显著的书写风格,包括楷书、行书及宫中奏折使用的半草书,但标注数据有限。本文构建多专家系统,将迭代微调生成的检查点作为领域专家,利用轻量级页面级图像分类器按视觉风格分配页面;当检查点池中无合适专家时,额外训练新专家。在三个固定测试集上,路由系统对每种风格的表现与选定专家一致:楷书0.30%字符错误率,奏折1.57%,行书4.83%。路由器页面级领域识别准确率达99.3%,与理想标签匹配精度相同。其中两个被选专家未专门针对其目标领域训练,仅行书专家为该领域特训。本文公开评估协议、路由器设计及逐页预测结果,确保可复现性。
原文摘要 · Abstract (English)
Historical Manchu OCR must accommodate various visually distinct writing styles, including regular script, running script, and the semi-cursive chancery hand used in palace memorials, despite limited labeled data. We study a multi-expert system that reuses checkpoints from an iterative fine-tuning process as domain specialists and uses a lightweight page-level image classifier to dispatch pages by visual style. When the checkpoint pool lacks a suitable specialist, we train an additional expert for that domain. On three frozen test sets, the routed system matches the selected specialist for each style at two-decimal precision: 0.30 percent CER on regular script, 1.57 percent on memorials, and 4.83 percent on running script. The router achieves 99.3 percent page-level domain accuracy and matches the domain-label oracle at the same precision. Two of the three selected specialists were not trained specifically for their final domain; only the running-script expert was trained with that domain as its target. We report the evaluation protocol, router design, and per-page predictions to make the comparison reproducible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。