arXiv:2510.01651cs.CV2025-10被引 1

用分层专家适配器提升青铜铭文识别,解决模糊、多样和罕见字难题

LadderMoE: Ladder-Side Mixture of Experts Adapters for Bronze Inscription Recognition

  • 采用分层MoE适配器增强CLIP模型,实现跨域动态专家分工
  • 在2.2万张图、近20万字符上验证,尾部类别准确率显著提升
  • 适合考古、古文字研究者,推动自动化铭文分析发展

青铜铭文(BI)刻于礼器之上,是早期汉字的重要阶段,对考古与历史研究至关重要。但自动识别面临严重视觉退化、图像/拓片/摹本等多源数据差异大、以及极长尾字符分布等挑战。为此,我们构建了大规模BI数据集,包含22454张全页图像和198598个标注字符,覆盖6658个唯一字类,支持跨域评估。基于此,提出两阶段检测-识别流水线:先定位铭文区域,再逐字转录。为应对异构域与稀有字,引入LadderMoE,通过分层式Mixture of Experts适配器增强预训练CLIP编码器,实现动态专家专精与更强鲁棒性。单字与全页识别任务的综合实验表明,该方法显著优于当前最优场景文本识别基线,在头部、中部、尾部类别及所有采集模态下均取得更优精度。成果为青铜铭文识别及下游考古分析奠定坚实基础。

原文摘要 · Abstract (English)

Bronze inscriptions (BI), engraved on ritual vessels, constitute a crucial stage of early Chinese writing and provide indispensable evidence for archaeological and historical studies. However, automatic BI recognition remains difficult due to severe visual degradation, multi-domain variability across photographs, rubbings, and tracings, and an extremely long-tailed character distribution. To address these challenges, we curate a large-scale BI dataset comprising 22454 full-page images and 198598 annotated characters spanning 6658 unique categories, enabling robust cross-domain evaluation. Building on this resource, we develop a two-stage detection-recognition pipeline that first localizes inscriptions and then transcribes individual characters. To handle heterogeneous domains and rare classes, we equip the pipeline with LadderMoE, which augments a pretrained CLIP encoder with ladder-style MoE adapters, enabling dynamic expert specialization and stronger robustness. Comprehensive experiments on single-character and full-page recognition tasks demonstrate that our method substantially outperforms state-of-the-art scene text recognition baselines, achieving superior accuracy across head, mid, and tail categories as well as all acquisition modalities. These results establish a strong foundation for bronze inscription recognition and downstream archaeological analysis.

青铜铭文多专家模型长尾识别考古视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。