构建首个韩文视觉文档检索基准,解决多页复杂文档的跨页信息整合难题。
KoViDoRe: Korean Visual Document Retrieval

- 基于真实韩文文档设计多阶段数据清洗流程,涵盖结构解析与人工验证
- 发现现有模型在处理表格、图表等结构化内容时表现不佳,跨页检索准确率显著下降
- 开源大规模训练数据集,助力开发专精于韩文视觉文档的检索模型
近年来多模态检索技术提升了从图文丰富文档(如PDF、报告)中获取信息的能力。然而,现有基准主要聚焦英文,对韩文复杂版式文档覆盖不足。多数韩文资源仅评估单页检索,难以反映需跨页证据聚合的真实场景。为此,我们提出KoViDoRe,一个面向韩文视觉文档检索的基准数据集。数据源自公开可得的韩文文档,包含表格、图示及多栏布局等多样结构。通过多阶段数据清洗流程——包括结构化文档解析、基于摘要与上下文的合成查询生成,以及人工验证的相关性标注——构建高质量数据集。在该基准上评估多种多模态检索模型,发现当前模型在处理结构化内容和多样化查询时表现有限。基于此,我们进一步构建大规模训练数据集Ko-VDR Train Public,以支持专用于韩文视觉文档的检索模型研发。KoViDoRe与Ko-VDR Train Public共同构成韩文视觉文档检索的统一评测与训练资源。
原文摘要 · Abstract (English)
Recent advances in multimodal retrieval have improved the ability to retrieve information from visually rich documents such as PDFs and reports. However, existing benchmarks remain largely centered on English and provide limited coverage of Korean visual documents with complex structures. Furthermore, most existing Korean resources primarily evaluate single-page retrieval, failing to capture realistic scenarios that require evidence aggregation across multiple pages. To address these gaps, we introduce KoViDoRe, a benchmark for Korean visual document retrieval. The dataset is constructed from publicly available Korean documents with diverse layouts, including tables, figures, and multi-column structures. We develop a multi-stage data curation pipeline consisting of structured document parsing, synthetic query generation using both summary-based and context-based strategies, and relevance mapping with human verification. Using KoViDoRe, we evaluate a wide range of multimodal retrieval models and observe that current models struggle to effectively handle Korean visual document retrieval, particularly in settings involving structured content and diverse query types. Motivated by this finding, we further curate a large-scale training dataset, Ko-VDR Train Public, to support the development of retrieval models tailored to Korean visual documents. Together, KoViDoRe and Ko-VDR Train Public provide a unified benchmark and training resource for Korean visual document retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。