arXiv:2512.18004cs.CVcs.AI2025-12中稿 · AAAI被引 3

用视觉大模型直接翻译手写法律文档,提升低资源语言处理效率

Seeing Justice Clearly: Handwritten Legal Document Translation with OCR and Vision-Language Models

  • 采用视觉大模型端到端直接翻译手写图像,省去传统分步流程
  • 在马拉地语手写法律文档上表现优于传统OCR+机器翻译流水线
  • 适合需要快速数字化法律文件的司法机构与非母语使用者

手写文本识别(HTR)与机器翻译在低资源语言如马拉地语中仍面临巨大挑战,因缺乏大规模数字化语料且手写风格差异显著。传统方法采用两阶段流水线:先通过OCR提取图像中的文本,再使用机器翻译模型进行翻译。本文探索并比较了传统OCR-MT流水线与统一视觉大语言模型在端到端直接翻译手写文本图像方面的性能。研究动机源于印度各级法院亟需可扩展、高精度的系统来数字化犯罪报告(FIR)、起诉书和证人陈述等法律记录。我们在一个精心构建的手写马拉地语法律文档数据集上评估两种方法,目标是实现即使在低资源环境下也能高效处理法律文件。结果提供了可操作的见解,有助于构建鲁棒、适合边缘部署的解决方案,提升非母语者及法律从业者的法律信息获取能力。

原文摘要 · Abstract (English)

Handwritten text recognition (HTR) and machine translation continue to pose significant challenges, particularly for low-resource languages like Marathi, which lack large digitized corpora and exhibit high variability in handwriting styles. The conventional approach to address this involves a two-stage pipeline: an OCR system extracts text from handwritten images, which is then translated into the target language using a machine translation model. In this work, we explore and compare the performance of traditional OCR-MT pipelines with Vision Large Language Models that aim to unify these stages and directly translate handwritten text images in a single, end-to-end step. Our motivation is grounded in the urgent need for scalable, accurate translation systems to digitize legal records such as FIRs, charge sheets, and witness statements in India's district and high courts. We evaluate both approaches on a curated dataset of handwritten Marathi legal documents, with the goal of enabling efficient legal document processing, even in low-resource environments. Our findings offer actionable insights toward building robust, edge-deployable solutions that enhance access to legal information for non-native speakers and legal professionals alike.

手写识别法律AI多模态翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。