arXiv:2512.19360cs.IR2025-12

用生成式向量搜索提升病理多模态模型的检索能力

Generative vector search to improve pathology foundation models across multimodal vision-language tasks

  • 通过条件采样生成查询相关的嵌入向量,实现更精准的跨模态检索
  • 在文献、病历和组织图像任务中提升10%-30%的检索准确率
  • 适合需要高精度多模态信息检索的研究者,尤其医学领域

检索增强生成通过引入外部知识源来减少大语言模型的幻觉并解决知识截止问题。然而,传统基于嵌入的检索难以捕捉多概念查询的复杂性,尤其是在生物医学领域,因为生物数据本身具有高维特性。例如,组学数据与临床报告同时包含大量分子、细胞和生理特征。我们提出随机潜在匹配(STHLM),一种生成式向量搜索方法,通过从文本或图像输入中采样条件嵌入来提升检索性能。类似于链式思维推理让语言模型‘思考更久’,STHLM使检索系统能‘搜索更广’,通过迭代采样实现。STHLM在多个基准测试中显著优于经典向量检索,涵盖科学文献、临床记录和组织图像,在测试时计算下提升10%-30%的检索性能,同时支持高达10倍的嵌入维度压缩。

原文摘要 · Abstract (English)

Retrieval-augmented generation improves large language models by grounding outputs in external knowledge sources, reducing hallucinations and addressing knowledge cutoffs. However, standard embedding-based retrieval fails to capture the complexity of multi-concept queries, particularly in domains like biomedicine, where biological data are inherently high-dimensional. For example,omics datasets, and clinical reports simultaneously exhibit numerous molecular, cellular, and physiological features. We present Stochastic Latent Matching (STHLM), a generative vector search method that samples query-conditioned embeddings from text or image inputs to enhance retrieval performance. Analogous to how Chain-of-Thought reasoning enables language models to "think longer" on complex problems, STHLM allows retrieval systems to "search wider" through iterative sampling. STHLM demonstrates critical improvements over classical vector retrieval across diverse benchmarks, including scientific literature, clinical notes, and tissue images, boosting retrieval performance by 10-30% through test-time compute (trading latency for accuracy), while enabling up to a 10-fold compression of embedding dimensions.

多模态检索生成式搜索病理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。