首个食品包装OCR多语言基准,解决标签文字密集、小字号难题
HalalBench: A Multilingual OCR Benchmark for Food Packaging Ingredient Extraction

- 构建1043张真实与合成图像数据集,覆盖14种语言
- 现有引擎在日文上全失效(F1=0.000),改进后提升36%准确率
- 适用于自动化清真认证系统研发,尤其关注多语言场景
目前缺乏针对食品包装的标准化OCR评估基准,尽管其在自动化清真食品验证中至关重要。现有基准侧重文档或场景文本,未涵盖成分标签的独特挑战:曲面、密集多语言文本及小于8pt字体。本文提出首个开源多语言食品包装OCR基准HalalBench,包含1,043张图像(50张真实,993张合成)和36,438个标注,以COCO格式呈现,覆盖14种语言。我们评估了docTR(F1=0.193)、ML Kit(0.180)、EasyOCR(0.167),但所有模型在日文上表现失败(F1=0.000)。聚类消融实验表明,后处理算法可提升36%的F1分数。研究结果通过生产级清真扫描工具HalalLens(https://halallens.no)验证,服务20多个国家。数据集与代码已开源。
原文摘要 · Abstract (English)
No standardized benchmark exists for evaluating OCR on food packaging, despite its critical role in automated halal food verification. Existing benchmarks target documents or scene text, missing the unique challenges of ingredient labels: curved surfaces, dense multilingual text, and sub-8pt fonts. We present HalalBench, the first open multilingual benchmark for food packaging OCR, comprising 1,043 images (50 real, 993 synthetic) with 36,438 annotations in COCO format spanning 14 languages. We evaluate four engines: docTR achieves F1=0.193, ML Kit 0.180, EasyOCR 0.167, while all fail on Japanese (F1=0.000). A clustering ablation shows 36% F1 improvement from our post-processing algorithm. We validate findings through HalalLens (https://halallens.no), a production halal scanner serving 20+ countries. Dataset and code are released under open licenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。