构建首个日文场景文本基准,专测视觉语言模型在真实环境中的理解能力。
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
- 新采集3241个真实图像,含112万字、3643种独特字符,覆盖竖排、混写等难点
- 设计三项任务:密集文本问答、收据关键信息提取、手写体识别,全面评估能力
- 发现模型对汉字识别仍存瓶颈,最佳模型平均得分仅0.64,适合日文多模态研究者
日文场景文本因其混合书写、频繁竖排及远超拉丁字母的字符数量,常被多语言基准忽略。现有日文视觉文本数据集多集中于扫描文档,真实场景文本仍待探索。为此,我们提出JaWildText,一个用于评估视觉语言模型(VLMs)在日文场景文本理解上的诊断性基准。该数据集包含3,241个实例,来自2,961张在日本实地拍摄的新图像,共标注112万字符,涵盖3,643种唯一字符类型。包含三个互补任务:(i) 密集场景文本视觉问答(STVQA),需基于多段文本证据推理;(ii) 收据关键信息抽取(KIE),测试移动端拍摄收据的布局感知结构化提取;(iii) 手写体OCR,评估跨媒体与书写方向的页面级转录能力。我们评测了14个开源权重的VLM,最优模型在三项任务上平均得分为0.64。错误分析表明,识别仍是主要瓶颈,尤其在汉字方面。JaWildText支持细粒度、脚本感知的诊断,将随评估代码一同发布。
原文摘要 · Abstract (English)
Japanese scene text poses challenges that multilingual benchmarks often fail to capture, including mixed scripts, frequent vertical writing, and a character inventory far larger than the Latin alphabet. Although Japanese is included in several multilingual benchmarks, these resources do not adequately capture the language-specific complexities. Meanwhile, existing Japanese visual text datasets have primarily focused on scanned documents, leaving in-the-wild scene text underexplored. To fill this gap, we introduce JaWildText, a diagnostic benchmark for evaluating vision-language models (VLMs) on Japanese scene text understanding. JaWildText contains 3,241 instances from 2,961 images newly captured in Japan, with 1.12 million annotated characters spanning 3,643 unique character types. It comprises three complementary tasks that vary in visual organization, output format, and writing style: (i) Dense Scene Text Visual Question Answering (STVQA), which requires reasoning over multiple pieces of visual text evidence; (ii) Receipt Key Information Extraction (KIE), which tests layout-aware structured extraction from mobile-captured receipts; and (iii) Handwriting OCR, which evaluates page-level transcription across various media and writing directions. We evaluate 14 open-weight VLMs and find that the best model achieves an average score of 0.64 across the three tasks. Error analyses show recognition remains the dominant bottleneck, especially for kanji. JaWildText enables fine-grained, script-aware diagnosis of Japanese scene text capabilities, and will be released with evaluation code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。