arXiv:2511.04910cs.CL2025-11被引 4

首个面向韩语官方文档的视觉文档检索基准数据集,解决多模态理解难题。

SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents

  • 基于361份真实韩语文档构建,涵盖表格、图表等复杂排版
  • 600组人工验证的查询-页面-答案三元组,覆盖六大公共领域
  • 支持文本与多模态检索双任务,揭示当前模型在跨模态推理中的不足

现有视觉文档检索(VDR)基准普遍忽略非英语语言及官方文件的结构复杂性。为填补这一空白,我们推出SDS KoPub VDR,首个大规模公开的韩语官方文档检索基准。数据集基于361份真实文档,包括256份KOGL Type 1许可文件和105份官方法律门户资料,包含表格、图表、多栏布局等复杂视觉元素。为确保评估可靠性,我们构建了600个经过人工验证的查询-页面-答案三元组,初始由多模态模型(如GPT-4o)生成,再经人工校验以保证事实准确性和上下文相关性。查询覆盖六个主要公共领域,按所需推理模式分为文本型、视觉型和跨模态型。我们在两个互补任务上评估:(1)纯文本检索,(2)融合视觉特征的多模态检索。双任务评估揭示显著性能差距,尤其在需要跨模态推理的多模态场景中,即使最先进模型也表现不佳。该基准为真实文档智能中的多模态AI发展提供可靠评估工具与路线图。数据集已开源:https://huggingface.co/datasets/SamsungSDS-Research/SDS-KoPub-VDR-Benchmark。

原文摘要 · Abstract (English)

Existing benchmarks for visual document retrieval (VDR) largely overlook non-English languages and the structural complexity of official publications. To address this gap, we introduce SDS KoPub VDR, the first large-scale, public benchmark for retrieving and understanding Korean public documents. The benchmark is built upon 361 real-world documents, including 256 files under the KOGL Type 1 license and 105 from official legal portals, capturing complex visual elements like tables, charts, and multi-column layouts. To establish a reliable evaluation set, we constructed 600 query-page-answer triples. These were initially generated using multimodal models (e.g., GPT-4o) and subsequently underwent human verification to ensure factual accuracy and contextual relevance. The queries span six major public domains and are categorized by the reasoning modality required: text-based, visual-based, and cross-modal. We evaluate SDS KoPub VDR on two complementary tasks: (1) text-only retrieval and (2) multimodal retrieval, which leverages visual features alongside text. This dual-task evaluation reveals substantial performance gaps, particularly in multimodal scenarios requiring cross-modal reasoning, even for state-of-the-art models. As a foundational resource, SDS KoPub VDR enables rigorous and fine-grained evaluation and provides a roadmap for advancing multimodal AI in real-world document intelligence. The dataset is available at https://huggingface.co/datasets/SamsungSDS-Research/SDS-KoPub-VDR-Benchmark.

文档检索多模态韩语基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。