对比轻量OCR与多模态模型在户外广告识别中的表现
Seeing the Signs: A Survey of Edge-Deployable OCR Models for Billboard Visibility Analysis
- 用合成天气噪声增强数据集,测试模型在真实场景下的文字识别能力
- 轻量CNN模型在裁剪文本上精度接近多模态模型,但计算成本低得多
- 适合边缘部署的广告可见性分析研究者参考
户外广告仍是现代营销的关键媒介,但真实条件下准确验证广告文字可见性仍具挑战。传统OCR流水线在裁剪文本识别上表现良好,但在复杂户外场景、字体变化和天气引起的视觉噪声下表现不佳。近期多模态视觉语言模型(VLMs)成为有前景的替代方案,可端到端理解场景而无需显式检测。本文系统评估了代表性VLMs(Qwen 2.5 VL 3B、InternVL3、SmolVLM2)与轻量级CNN基线(PaddleOCRv4)在两个公开数据集(ICDAR 2015和SVT)上的表现,数据集通过合成天气失真进行增强以模拟真实退化。结果表明,尽管选定VLM在整体场景理解上更优,但轻量CNN模型在裁剪文本识别上仍能实现具有竞争力的准确率,且计算开销仅为前者的一小部分——这对边缘部署至关重要。为促进后续研究,我们公开了带天气增强的基准数据集与评估代码。
原文摘要 · Abstract (English)
Outdoor advertisements remain a critical medium for modern marketing, yet accurately verifying billboard text visibility under real-world conditions is still challenging. Traditional Optical Character Recognition (OCR) pipelines excel at cropped text recognition but often struggle with complex outdoor scenes, varying fonts, and weather-induced visual noise. Recently, multimodal Vision-Language Models (VLMs) have emerged as promising alternatives, offering end-to-end scene understanding with no explicit detection step. This work systematically benchmarks representative VLMs - including Qwen 2.5 VL 3B, InternVL3, and SmolVLM2 - against a compact CNN-based OCR baseline (PaddleOCRv4) across two public datasets (ICDAR 2015 and SVT), augmented with synthetic weather distortions to simulate realistic degradation. Our results reveal that while selected VLMs excel at holistic scene reasoning, lightweight CNN pipelines still achieve competitive accuracy for cropped text at a fraction of the computational cost-an important consideration for edge deployment. To foster future research, we release our weather-augmented benchmark and evaluation code publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。