arXiv:2604.23685cs.CV2026-04

解决暗光场景文字识别难题,构建新数据集并提出联合优化方法

Reading in the Dark: Low-light Scene Text Recognition

论文配图:Reading in the Dark: Low-light Scene Text Recognition
图 1 · 摘自论文原文
  • 构建1.1万张合成暗光图像数据集,融合真实夜景样本
  • 联合训练增强模块与OCR模型,显著提升暗光识别准确率
  • 揭示光照强度阈值对识别效果的关键影响,适合自动驾驶等应用

暗光环境下精准文字识别对自动驾驶、智能监控等智能系统至关重要,但光照不足和噪声干扰问题尚未充分研究。为此,我们构建了LSTR数据集,包含从ICDAR2015、IIIT5K和WordArt等清晰图像生成的11,273张暗光图像,并创建了包含60张真实夜间街景图像的ESTR数据集,用于专项评估。提出两种解决方案:一是对OCR模型进行微调及LoRA微调;二是将低光图像增强(LLIE)模块与OCR模型联合训练。特别提出一种重渲染式低光增强(RLLIE)模块,在真实数据上表现更优。实验表明,单一增强或识别模型在暗光下性能不佳,而专为文字设计的联合训练策略更具优势。研究还回答了核心问题:‘亮度达到何种程度才足够支持有效文字识别?’ 提供了全面基准以推动后续鲁棒暗光文字识别研究。

原文摘要 · Abstract (English)

Accurate text recognition in low-light environments is essential for intelligent systems in applications ranging from autonomous vehicles to smart surveillance. However, challenges such as poor illumination and noise interference remain underexplored. To address this gap, we introduce LSTR, a large-scale Low-light Scene Text Recognition dataset comprising 11,273 low-light images generated from well-lit datasets (ICDAR2015, IIIT5K, and WordArt), along with ESTR, which includes 60 real nighttime street-scene images in English and Spanish for exclusive evaluation. We explore two solution strategies: (1) employing Optical Character Recognition (OCR) models with fine-tuning and LoRA-based fine-tuning and (2) a joint training strategy that integrates a low-light image enhancement (LLIE) module with an OCR model. In particular, we propose a novel re-render LLIE (RLLIE) module, which demonstrates improved performance on real-world data. Through extensive experimentation, we analyze various training strategies and address a key research question: \emph{How bright is bright enough for effective scene text recognition?} Our results indicate that standalone LLIE or OCR models perform inadequately under low-light conditions, highlighting the advantages of specialized, jointly trained text-centric approaches. Additionally, we provide a comprehensive benchmark to support future research in robust low-light scene text recognition. https://huggingface.co/datasets/lumimusta/Low-light_Scene_Text_Dataset.

文字识别暗光处理联合训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。