首个面向现代葡萄牙语的视觉文本提取基准,填补OCR空白。
PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction

- 构建首个面向当代葡语的视觉文本提取基准
- 真实场景下模型性能显著下降,合成数据效果不佳
- 适合多语言OCR研究者与本地化应用开发者
欧洲葡萄牙语(pt-PT)在光学字符识别(OCR)基准中几乎缺失,现有基准多聚焦历史文献。本文针对现代应用场景,提出PorTEXTO——首个面向当代、文化相关的pt-PT视觉文本提取基准。为保证质量,采用前沿大视觉语言模型(LVLM)生成转录,并经母语者逐项审校。实验发现,多数模型在从合成数据到真实样本时性能骤降;当前,专用多语言数据比模型规模或分辨率更能提升pt-PT表现,据此释放开放的pt-PT OCR资源。
原文摘要 · Abstract (English)
European Portuguese (pt-PT) is largely absent from Optical Character Recognition (OCR) benchmarks, which skew toward high-resource languages. The few benchmarks that cover pt-PT focus on historical artifacts and literature. This work addresses modern OCR applications, introducing PorTEXTO, the first benchmark for contemporary and culturally relevant pt-PT visual text extraction. To ascertain quality, we employ an annotation pipeline combining transcriptions from a frontier LVLM with exhaustive review by native speakers. We observe a sharp performance drop from synthetic to real world samples in most models, and find that, currently, specialized multilingual data is a better driver for pt-PT performance than model size or resolution budget, motivating the release of open pt-PT OCR resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。