针对石刻文字二值化难题,提出自适应分块策略提升识别精度。
Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization
- 基于注意力U-Net与动态分块,聚焦细微笔画结构
- 在多类石刻上实现显著性能提升,零样本泛化至非印欧语系
- 适合历史文献数字化、古文字OCR研究者使用
二值化是历史文物文本提取的关键第一步。石刻图像因字符与背景对比度差、表面不均匀退化、干扰杂质及文字密度和布局高度可变,给二值化带来严峻挑战,导致现有方法常失效且难以分离连贯字符区域。许多方法通过分块提升文本片段分辨率以改善性能。本文提出一种鲁棒自适应的分块策略,用于处理具有挑战性的印地语系石刻。利用该策略生成的分块训练注意力U-Net进行二值化,注意力机制使模型聚焦于细微结构线索,动态采样与分块选择方法确保模型能克服表面噪声和布局不规则性。我们还构建了一个精细标注的印地语系石刻字符片段级数据集。实验表明,该新型分块机制显著提升经典与深度学习基线的二值化性能。尽管仅在单一印地语文本数据集上训练,模型仍展现出对其他印地语及非印地语文字的强零样本泛化能力,凸显其鲁棒性与跨脚本通用性。该方法生成的清晰结构化文本表示,为后续脚本识别、OCR及历史文本分析奠定基础。项目页面:https://ihdia.iiit.ac.in/shilalekhya-binarization/
原文摘要 · Abstract (English)
Binarization is a popular first step towards text extraction in historical artifacts. Stone inscription images pose severe challenges for binarization due to poor contrast between etched characters and the stone background, non-uniform surface degradation, distracting artifacts, and highly variable text density and layouts. These conditions frequently cause existing binarization techniques to fail and struggle to isolate coherent character regions. Many approaches sub-divide the image into patches to improve text fragment resolution and improve binarization performance. With this in mind, we present a robust and adaptive patching strategy to binarize challenging Indic inscriptions. The patches from our approach are used to train an Attention U-Net for binarization. The attention mechanism allows the model to focus on subtle structural cues, while our dynamic sampling and patch selection method ensures that the model learns to overcome surface noise and layout irregularities. We also introduce a carefully annotated, pixel-precise dataset of Indic stone inscriptions at the character-fragment level. We demonstrate that our novel patching mechanism significantly boosts binarization performance across classical and deep learning baselines. Despite training only on single script Indic dataset, our model exhibits strong zero-shot generalization to other Indic and non-indic scripts, highlighting its robustness and script-agnostic generalization capabilities. By producing clean, structured representations of inscription content, our method lays the foundation for downstream tasks such as script identification, OCR, and historical text analysis. Project page: https://ihdia.iiit.ac.in/shilalekhya-binarization/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。