统一视觉与文本特征,实现高效文档检索。
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval

- 融合视觉与文本特征,构建鲁棒文档表示。
- 通过知识蒸馏提升纯视觉模型的语义理解能力。
- 适合需要高精度、无解析文档检索的场景。
真实场景下的文档检索面临格式与模态多样性挑战。传统文本方法依赖定制解析,忽略版式信息且易出错;近期无解析视觉方法在文本密集场景中难以捕捉细粒度语义。为此,我们提出 extbf{Unveil},一种统一的视觉-文本嵌入框架,有效整合文本与视觉特征以实现稳健文档表征。通过知识蒸馏,将视觉-文本嵌入模型的语义理解能力迁移至纯视觉模型,实现高效无解析检索的同时保持语义保真。实验表明,该方法显著优于现有技术,知识蒸馏成功缩小了视觉-文本与纯视觉方法间的性能差距,提升了检索准确率与效率。
原文摘要 · Abstract (English)
Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities. Traditional text-based approaches rely on tailored parsing techniques that disregard layout information and are prone to errors, while recent parsing-free visual methods often struggle to capture fine-grained textual semantics in text-rich scenarios. To address these limitations, we propose \textbf{Unveil}, a novel visual-textual embedding framework that effectively integrates textual and visual features for robust document representation. Through knowledge distillation, we transfer the semantic understanding capabilities from the visual-textual embedding model to a purely visual model, enabling efficient parsing-free retrieval while preserving semantic fidelity. Experimental results demonstrate that our visual-textual embedding method surpasses existing approaches, while knowledge distillation successfully bridges the performance gap between visual-textual and visual-only methods, improving both retrieval accuracy and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。