arXiv:2509.15211cs.CL2025-09

比较多种幻灯片检索方法,找出高效实用的方案。

What's the Best Way to Retrieve Slides? A Comparative Study of Multimodal, Caption-Based, and Hybrid Retrieval Techniques

  • 用视觉语言模型生成幻灯片描述,降低存储需求
  • 混合检索比纯文本或纯视觉方法效果更好
  • 兼顾速度、存储与精度,适合实际应用

幻灯片作为连接演示文稿与书面文档的数字报告,在学术和企业场景中广泛使用。其融合文字、图像和图表的多模态特性,给检索增强生成系统带来挑战,检索质量直接影响下游表现。传统方法通常对不同模态分别索引,增加复杂性并损失上下文信息。本文对比了多种幻灯片检索技术,包括基于视觉的晚交互嵌入模型(如 ColPali)、视觉重排序器、以及结合密集检索与 BM25 的混合检索策略,并引入文本重排序和互惠排名融合等优化方法。同时评估了一种基于视觉-语言模型的全新图文生成流水线,结果显示其在保持相近检索性能的同时,显著降低嵌入存储开销。研究还综合分析了各方法的运行时性能与存储成本,为真实场景下高效、稳健的幻灯片检索系统选型与开发提供实用指导。

原文摘要 · Abstract (English)

Slide decks, serving as digital reports that bridge the gap between presentation slides and written documents, are a prevalent medium for conveying information in both academic and corporate settings. Their multimodal nature, combining text, images, and charts, presents challenges for retrieval-augmented generation systems, where the quality of retrieval directly impacts downstream performance. Traditional approaches to slide retrieval often involve separate indexing of modalities, which can increase complexity and lose contextual information. This paper investigates various methodologies for effective slide retrieval, including visual late-interaction embedding models like ColPali, the use of visual rerankers, and hybrid retrieval techniques that combine dense retrieval with BM25, further enhanced by textual rerankers and fusion methods like Reciprocal Rank Fusion. A novel Vision-Language Models-based captioning pipeline is also evaluated, demonstrating significantly reduced embedding storage requirements compared to visual late-interaction techniques, alongside comparable retrieval performance. Our analysis extends to the practical aspects of these methods, evaluating their runtime performance and storage demands alongside retrieval efficacy, thus offering practical guidance for the selection and development of efficient and robust slide retrieval systems for real-world applications.

幻灯片检索多模态混合检索视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。