arXiv:2508.19944cs.CVcs.CL2025-08EMNLP被引 5

为韩语文本丰富视觉问答构建新基准,助力多语言模型评估

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts

  • 构建韩语图文理解与推理的综合评测基准
  • 覆盖15个领域26种图像类型,支持多维度评估
  • 提供可复用的自动化数据生成流程,适合多语言研究

视觉语言模型在复杂多样的真实场景中理解与推理文本面临巨大挑战。尽管英语等高资源语言已有丰富的文本丰富型视觉问答(VQA)数据集,但韩语等低资源语言仍缺乏全面的评测基准,制约了模型的评估与比较。为此,本文提出KRETA——一个专为韩语文本丰富型VQA设计的基准,涵盖15个领域和26种图像类型,可深入评估模型的视觉文本理解与推理能力。我们还设计了一套半自动化的VQA生成流程,结合精细化的图像分步分解与严格的七项评估指标,保障数据质量。该框架具备可扩展性,可推广至其他语言,推动多语言视觉语言模型研究。代码与数据集已开源。

原文摘要 · Abstract (English)

Understanding and reasoning over text within visual contexts poses a significant challenge for Vision-Language Models (VLMs), given the complexity and diversity of real-world scenarios. To address this challenge, text-rich Visual Question Answering (VQA) datasets and benchmarks have emerged for high-resource languages like English. However, a critical gap persists for low-resource languages such as Korean, where the lack of comprehensive benchmarks hinders robust model evaluation and comparison. To bridge this gap, we introduce KRETA, a benchmark for Korean Reading and rEasoning in Text-rich VQA Attuned to diverse visual contexts. KRETA facilitates an in-depth evaluation of both visual text understanding and reasoning capabilities, while also supporting a multifaceted assessment across 15 domains and 26 image types. Additionally, we introduce a semi-automated VQA generation pipeline specifically optimized for text-rich settings, leveraging refined stepwise image decomposition and a rigorous seven-metric evaluation protocol to ensure data quality. While KRETA is tailored for Korean, we hope our adaptable and extensible pipeline will facilitate the development of similar benchmarks in other languages, thereby accelerating multilingual VLM research. The code and dataset for KRETA are available at https://github.com/tabtoyou/KRETA.

视觉问答多语言韩语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。