构建菜单图文翻译评测集,精准评估大模型对复杂排版的理解与翻译能力。
Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models
- 设计包含中英菜单的复杂排版数据集,涵盖字体多样性和文化元素。
- 自动评估结果与专业人工评价高度一致,验证了评测框架可靠性。
- 适合研究视觉语言模型在跨文化交流中的应用与优化方向。
大型视觉语言模型(LVLMs)在文档理解领域发展迅速,尤其在光学字符识别(OCR)和多语言翻译方面表现突出。然而,现有评估体系如OCRBench主要关注短文本和简单布局下的准确性,对复杂排版长文本的理解能力评估严重不足。为此,本文提出菜单OCR与翻译评测基准MOTBench,聚焦菜单翻译在跨文化交流中的关键作用。该基准要求模型准确识别菜单中每道菜品及其价格、单位信息,全面评估其视觉理解与语言处理能力。数据集包含中英文菜单,具有复杂布局、多样字体及文化特异性元素,并配有精确的人工标注。实验表明,自动评估结果与专业人工评价高度一致。我们评估了多种公开的先进LVLMs,通过分析输出揭示其优劣势,为未来模型改进提供重要参考。MOTBench已开源:https://github.com/gitwzl/MOTBench。
原文摘要 · Abstract (English)
The rapid advancement of large vision-language models (LVLMs) has significantly propelled applications in document understanding, particularly in optical character recognition (OCR) and multilingual translation. However, current evaluations of LVLMs, like the widely used OCRBench, mainly focus on verifying the correctness of their short-text responses and long-text responses with simple layout, while the evaluation of their ability to understand long texts with complex layout design is highly significant but largely overlooked. In this paper, we propose Menu OCR and Translation Benchmark (MOTBench), a specialized evaluation framework emphasizing the pivotal role of menu translation in cross-cultural communication. MOTBench requires LVLMs to accurately recognize and translate each dish, along with its price and unit items on a menu, providing a comprehensive assessment of their visual understanding and language processing capabilities. Our benchmark is comprised of a collection of Chinese and English menus, characterized by intricate layouts, a variety of fonts, and culturally specific elements across different languages, along with precise human annotations. Experiments show that our automatic evaluation results are highly consistent with professional human evaluation. We evaluate a range of publicly available state-of-the-art LVLMs, and through analyzing their output to identify the strengths and weaknesses in their performance, offering valuable insights to guide future advancements in LVLM development. MOTBench is available at https://github.com/gitwzl/MOTBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。