arXiv:2412.04137cs.CVcs.AI2024-12中稿 · ed被引 3

不依赖OCR,直接比对文档图像找文字变化。

Text Change Detection in Multilingual Documents Using Image Comparison

  • 用图像级对比替代OCR,实现跨语言文字变更检测。
  • 生成双向变化分割图,准确识别多语言文档差异区域。
  • 自建多语言印刷/扫描文档数据集,支持真实场景评估。

文档比对通常依赖光学字符识别(OCR)技术,但需为每份文档选择合适的语言模型,且多语言或混合语言模型性能有限。为克服此问题,我们提出一种针对多语言文档的文本变更检测(TCD)方法,采用专为图像比对设计的模型。与基于OCR的方法不同,本方法通过词级文本图像到图像的对比来检测变更,生成源文档与目标文档之间的双向变更分割图。为在无需显式文本对齐或缩放预处理的情况下提升性能,我们利用多尺度注意力特征间的相关性。同时构建了一个基准数据集,包含多种语言的实际印刷和扫描词对,用于评估模型。我们在自建数据集及公开基准数据集Distorted Document Images和LRDE Document Binarization Dataset上验证了该方法,并与最先进的语义分割与变更检测模型以及传统OCR模型进行比较。

原文摘要 · Abstract (English)

Document comparison typically relies on optical character recognition (OCR) as its core technology. However, OCR requires the selection of appropriate language models for each document and the performance of multilingual or hybrid models remains limited. To overcome these challenges, we propose text change detection (TCD) using an image comparison model tailored for multilingual documents. Unlike OCR-based approaches, our method employs word-level text image-to-image comparison to detect changes. Our model generates bidirectional change segmentation maps between the source and target documents. To enhance performance without requiring explicit text alignment or scaling preprocessing, we employ correlations among multi-scale attention features. We also construct a benchmark dataset comprising actual printed and scanned word pairs in various languages to evaluate our model. We validate our approach using our benchmark dataset and public benchmarks Distorted Document Images and the LRDE Document Binarization Dataset. We compare our model against state-of-the-art semantic segmentation and change detection models, as well as to conventional OCR-based models.

文档比对多语言图像对比文本检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。