arXiv:2508.17079cs.IRcs.AI2025-08EMNLP被引 7

用多模态大模型生成预提问,实现跨域跨语言文档检索

Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation

  • 用大模型生成跨模态预提问,提升检索匹配精度
  • 在跨域和多语言数据上超越现有方法,指标全面领先
  • 适合处理隐私文档和未知领域检索场景

多模态大语言模型的快速发展使信息检索突破纯文本限制,能够处理融合文字与图像的复杂现实文档。然而,多数文档属于个人或企业私有,现有检索器在面对未见领域或语言时表现不佳。为此,我们提出PREMIR框架,利用MLLM的广泛知识,在检索前生成跨模态预提问(preQs)。不同于传统在单一向量空间比较嵌入的方法,PREMIR通过多个互补模态的preQs将匹配范围扩展至词级别。实验表明,PREMIR在分布外基准上达到顶尖性能,涵盖闭域和多语言设置,在所有检索指标上均优于强基线。深入消融研究验证了各组件贡献,对生成preQs的定性分析进一步凸显模型在真实场景中的鲁棒性。

原文摘要 · Abstract (English)

Rapid advances in Multimodal Large Language Models (MLLMs) have expanded information retrieval beyond purely textual inputs, enabling retrieval from complex real world documents that combine text and visuals. However, most documents are private either owned by individuals or confined within corporate silos and current retrievers struggle when faced with unseen domains or languages. To address this gap, we introduce PREMIR, a simple yet effective framework that leverages the broad knowledge of an MLLM to generate cross modal pre questions (preQs) before retrieval. Unlike earlier multimodal retrievers that compare embeddings in a single vector space, PREMIR leverages preQs from multiple complementary modalities to expand the scope of matching to the token level. Experiments show that PREMIR achieves state of the art performance on out of distribution benchmarks, including closed domain and multilingual settings, outperforming strong baselines across all retrieval metrics. We confirm the contribution of each component through in depth ablation studies, and qualitative analyses of the generated preQs further highlight the model's robustness in real world settings.

多模态检索零样本大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。