首个直接将手写莫迪文转写为天城文的视觉语言模型
Historic Scripts to Modern Vision: A Novel Dataset and A VLM Framework for Transliteration of Modi Script to Devanagari
- 提出新视觉语言模型MoScNet,通过知识蒸馏提升转写性能
- 在2043张莫迪文图像上实现高精度转写,学生模型参数量仅为教师1/163
- 适合文化遗产数字化、历史文献研究者使用
中世纪印度用莫迪文书写马拉地语,现存约4000万份文献因保存状况差尚未转写。目前仅有少数专家能将该文转为英文或天城文。以往研究多聚焦单个字符识别。本文构建了包含2,043张莫迪文文档图像及其对应天城文转写文本的MoDeTrans数据集,并提出MoScNet(莫迪文网络)框架,一个用于直接将手写莫迪文图像转写为天城文的视觉-语言模型。该框架采用知识蒸馏技术,使学生模型在参数量仅为教师模型1/163的情况下,性能优于教师模型。本工作首次实现从手写莫迪文到天城文的端到端转写,且在光学字符识别(OCR)任务上表现良好。
原文摘要 · Abstract (English)
In medieval India, the Marathi language was written using the Modi script. The texts written in Modi script include extensive knowledge about medieval sciences, medicines, land records and authentic evidence about Indian history. Around 40 million documents are in poor condition and have not yet been transliterated. Furthermore, only a few experts in this domain can transliterate this script into English or Devanagari. Most of the past research predominantly focuses on individual character recognition. A system that can transliterate Modi script documents to Devanagari script is needed. We propose the MoDeTrans dataset, comprising 2,043 images of Modi script documents accompanied by their corresponding textual transliterations in Devanagari. We further introduce MoScNet (\textbf{Mo}di \textbf{Sc}ript \textbf{Net}work), a novel Vision-Language Model (VLM) framework for transliterating Modi script images into Devanagari text. MoScNet leverages Knowledge Distillation, where a student model learns from a teacher model to enhance transliteration performance. The final student model of MoScNet has better performance than the teacher model while having 163$\times$ lower parameters. Our work is the first to perform direct transliteration from the handwritten Modi script to the Devanagari script. MoScNet also shows competitive results on the optical character recognition (OCR) task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。