轻量级模型实现跨文字手写识别,精准且省资源。
Language-Agnostic Visual Embeddings for Cross-Script Handwriting Retrieval
- 用双编码器学习统一的、风格无关的视觉嵌入
- 在多种文字间检索准确率超越28个基线模型
- 适合资源受限设备部署,支持跨语言查询
手写文字检索对数字档案至关重要,但受书写差异大和跨语言语义鸿沟影响,仍具挑战。现有大规模视觉-语言模型虽有潜力,却因计算成本高难以在边缘设备实用。为此,我们提出一种轻量级非对称双编码框架,学习统一且风格不变的视觉嵌入。通过联合优化实例级对齐与类别级语义一致性,将视觉嵌入锚定于语言无关的语义原型,实现跨文字与书写风格的不变性。实验表明,该方法在同语言检索基准上优于28个基线,达到当前最优性能。进一步开展显式跨语言检索(查询语言与目标语言不同),验证了所学表示的跨语言有效性。仅需现有模型极小参数量即可实现高精度与低资源消耗,适用于高效跨文字手写检索。
原文摘要 · Abstract (English)
Handwritten word retrieval is vital for digital archives but remains challenging due to large handwriting variability and cross-lingual semantic gaps. While large vision-language models offer potential solutions, their prohibitive computational costs hinder practical edge deployment. To address this, we propose a lightweight asymmetric dual-encoder framework that learns unified, style-invariant visual embeddings. By jointly optimizing instance-level alignment and class-level semantic consistency, our approach anchors visual embeddings to language-agnostic semantic prototypes, enforcing invariance across scripts and writing styles. Experiments show that our method outperforms 28 baselines and achieves state-of-the-art accuracy on within-language retrieval benchmarks. We further conduct explicit cross-lingual retrieval, where the query language differs from the target language, to validate the effectiveness of the learned cross-lingual representations. Achieving strong performance with only a fraction of the parameters required by existing models, our framework enables accurate and resource-efficient cross-script handwriting retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。