arXiv:2605.24530cs.CLcs.CV2026-05ACL被引 8

统一视觉与文本特征,实现高效文档检索。

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval

论文配图:Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
图 1 · 摘自论文原文
  • 融合视觉与文本特征,构建鲁棒文档表示。
  • 通过知识蒸馏提升纯视觉模型的语义理解能力。
  • 适合需要高精度、无解析文档检索的场景。

真实场景下的文档检索面临格式与模态多样性挑战。传统文本方法依赖定制解析,忽略版式信息且易出错;近期无解析视觉方法在文本密集场景中难以捕捉细粒度语义。为此,我们提出 extbf{Unveil},一种统一的视觉-文本嵌入框架,有效整合文本与视觉特征以实现稳健文档表征。通过知识蒸馏,将视觉-文本嵌入模型的语义理解能力迁移至纯视觉模型,实现高效无解析检索的同时保持语义保真。实验表明,该方法显著优于现有技术,知识蒸馏成功缩小了视觉-文本与纯视觉方法间的性能差距,提升了检索准确率与效率。

原文摘要 · Abstract (English)

Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities. Traditional text-based approaches rely on tailored parsing techniques that disregard layout information and are prone to errors, while recent parsing-free visual methods often struggle to capture fine-grained textual semantics in text-rich scenarios. To address these limitations, we propose \textbf{Unveil}, a novel visual-textual embedding framework that effectively integrates textual and visual features for robust document representation. Through knowledge distillation, we transfer the semantic understanding capabilities from the visual-textual embedding model to a purely visual model, enabling efficient parsing-free retrieval while preserving semantic fidelity. Experimental results demonstrate that our visual-textual embedding method surpasses existing approaches, while knowledge distillation successfully bridges the performance gap between visual-textual and visual-only methods, improving both retrieval accuracy and efficiency.

文档检索多模态知识蒸馏视觉文本融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。