arXiv:2608.01389cs.AI2026-08

打造高效韩文文档检索模型,单向量+精心训练实现高性能

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval

  • 用单向量表示韩文文档,结合双语监督与难负样本挖掘
  • 20亿参数模型在韩文检索任务上超越80亿参数单向量和多向量基线
  • 适合需要轻量高效韩文视觉检索的场景,无需大模型或复杂结构

视觉文档检索(VDR)直接将文本查询与文档图像匹配,保留了文本提取中可能丢失的视觉与结构信息。然而现有VDR模型和训练资源仍以英语为主,且许多高性能系统依赖大型骨干网络或存储密集的多向量表示。为此,我们提出KoVRE:面向韩文视觉文档检索的嵌入模型,搭配一套完整的训练方法。模型在708,729对韩英查询-页面数据上进行训练,采用正例感知的难负样本挖掘,并对训练数据组成、难负样本处理及重排序知识蒸馏进行了受控分析。在多个韩文视觉文档检索基准上,我们的20亿参数模型显著优于基础骨干模型,不仅超过其80亿参数的单向量版本,还优于强基线多向量模型。结果表明,有针对性的双语监督与精心设计的训练策略,可在不依赖扩展骨干网络或多向量表示的前提下,有效构建跨多样文档领域的韩文VDR模型。

原文摘要 · Abstract (English)

Visual Document Retrieval (VDR) directly matches text queries against document images, preserving visual and structural information that may be lost during text extraction. However, existing VDR models and training resources remain predominantly English-centric, while many high-performing systems rely on massive backbones or storage-intensive multi-vector representations. To address these limitations, we introduce KoVRE: Korean Visual Document Retrieval Embedding, a single-vector retriever for Korean visual documents, alongside a comprehensive training recipe. We train the model on 708,729 Korean and English query-page pairs using positive-aware hard-negative mining and conduct controlled analyses of training-data composition, hard-negative treatment, and reranker-based knowledge distillation. Across Korean visual document retrieval benchmarks, our 2B model substantially improves over the base backbone model, outperforming both its 8B single-vector counterpart and a strong multi-vector baseline. These results demonstrate that targeted bilingual supervision and our carefully designed training strategies can produce a highly effective Korean VDR model across diverse document domains, without requiring a scaled-up backbone or multi-vector representations.

视觉检索韩文处理单向量轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。