用多模态大模型提升搜索理解与安全过滤能力
CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model
- 分阶段处理图像上下文与用户意图,生成精准查询
- 在知识问答与安全检测任务上超越现有模型表现
- 适合需要高精度多模态搜索的组织级应用
将检索增强生成(RAG)与多模态大语言模型(MLLMs)结合,已推动信息检索革新并拓展AI实际应用。然而,现有系统在准确理解用户意图、采用多样化检索策略及有效过滤不当响应方面仍存挑战,限制了其效果。本文提出一种新型多模态搜索框架CUE-M,通过多阶段流程实现图像上下文增强、意图优化、上下文查询生成、外部API集成与相关性过滤。CUE-M引入融合图像、文本和多模态分类器的强效过滤机制,动态适配组织政策定义的实例与类别级关注点。在真实数据集及公开基准测试(知识驱动VQA与安全性评估)上的大量实验表明,CUE-M优于基线模型,建立新SOTA,显著提升多模态检索系统能力。
原文摘要 · Abstract (English)
The integration of Retrieval-Augmented Generation (RAG) with Multimodal Large Language Models (MLLMs) has revolutionized information retrieval and expanded the practical applications of AI. However, current systems struggle in accurately interpreting user intent, employing diverse retrieval strategies, and effectively filtering unintended or inappropriate responses, limiting their effectiveness. This paper introduces Contextual Understanding and Enhanced Search with MLLM (CUE-M), a novel multimodal search framework that addresses these challenges through a multi-stage pipeline comprising image context enrichment, intent refinement, contextual query generation, external API integration, and relevance-based filtering. CUE-M incorporates a robust filtering pipeline combining image-based, text-based, and multimodal classifiers, dynamically adapting to instance- and category-specific concern defined by organizational policies. Extensive experiments on real-word datasets and public benchmarks on knowledge-based VQA and safety demonstrated that CUE-M outperforms baselines and establishes new state-of-the-art results, advancing the capabilities of multimodal retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。