arXiv:2502.16781cs.CL2025-02中稿 · CIKM 2025被引 5

测试多语言文字识别错误对大模型问答性能的影响

Evaluating Robustness of LLMs in Question Answering on Multilingual Noisy OCR Data

  • 构建含三语噪声的多语言问答数据集
  • 模型在噪声文本上表现显著下降
  • 揭示历史文档数字化中的鲁棒性短板

光学字符识别(OCR)在数字化历史多语言文献中至关重要,但文本提取中的错误(如字符插入、删除、替换)会严重影响下游问答任务。本文系统分析了OCR噪声对多语言问答系统的影响,构建了包含50,000个问答对的MultiOCR-QA数据集,涵盖英语、法语和德语,源自带不同级别与类型噪声的OCR文献。我们评估了多种先进大语言模型在不同错误条件下的表现,聚焦三大类主要OCR错误。结果表明,问答系统对OCR噪声极为敏感,在噪声文本上表现明显下降。通过对比清洁与噪声文本下的模型表现,揭示了现有方法的局限性,并强调在历史文献数字化中需发展更具噪声鲁棒性的问答系统。

原文摘要 · Abstract (English)

Optical Character Recognition (OCR) plays a crucial role in digitizing historical and multilingual documents, yet OCR errors - imperfect extraction of text, including character insertion, deletion, and substitution can significantly impact downstream tasks like question-answering (QA). In this work, we conduct a comprehensive analysis of how OCR-induced noise affects the performance of Multilingual QA Systems. To support this analysis, we introduce a multilingual QA dataset MultiOCR-QA, comprising 50K question-answer pairs across three languages, English, French, and German. The dataset is curated from OCR-ed historical documents, which include different levels and types of OCR noise. We then evaluate how different state-of-the-art Large Language Models (LLMs) perform under different error conditions, focusing on three major OCR error types. Our findings show that QA systems are highly prone to OCR-induced errors and perform poorly on noisy OCR text. By comparing model performance on clean versus noisy texts, we provide insights into the limitations of current approaches and emphasize the need for more noise-resilient QA systems in historical digitization contexts.

多语言OCR噪声问答系统历史文献

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。