arXiv:2509.08216cs.IR2025-09被引 2

用视觉语言模型提升计算机教材图文检索效果

Vector embedding of multi-modal texts: a tool for discovery?

  • 用VLM生成图文联合向量嵌入,存入向量库
  • 余弦相似度在75个自然语言查询中表现最佳
  • 适合数字图书馆构建智能发现系统

计算机科学文本富含叙述内容与图表、算法、图像、标注图示等多模态信息。本研究探索基于视觉语言模型(VLM)的向量多模态检索在跨图文内容发现中的潜力。基于3600余页数字化教材(主要来自计算机科学教材),利用VLM生成同时捕捉文本与视觉语义的多向量表示,并存入向量数据库。设计了75个自然语言查询基准,对比不同相似度度量(距离)下的检索性能与真实答案。结果表明,余弦相似度在语义与视觉相关性检索上最有效。进一步讨论了向量数据库与多模态嵌入在实际信息检索中的可行性。论文旨在为数字图书馆的智能发现提供设计启示。

原文摘要 · Abstract (English)

Computer science texts are particularly rich in both narrative content and illustrative charts, algorithms, images, annotated diagrams, etc. This study explores the extent to which vector-based multimodal retrieval, powered by vision-language models (VLMs), can improve discovery across multi-modal (text and images) content. Using over 3,600 digitized textbook pages largely from computer science textbooks and a Vision Language Model (VLM), we generate multi-vector representations capturing both textual and visual semantics. These embeddings are stored in a vector database. We issue a benchmark of 75 natural language queries and compare retrieval performance to ground truth and across four similarity (distance) measures. The study is intended to expose both the strengths and weakenesses of such an approach. We find that cosine similarity most effectively retrieves semantically and visually relevant pages. We further discuss the practicality of using a vector database and multi-modal embedding for operational information retrieval. Our paper is intended to offer design insights for discovery over digital libraries. Keywords: Vector embedding, multi-modal document retrieval, vector database benchmark, digital library discovery

多模态检索向量数据库数字图书馆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。