arXiv:2501.15470cs.IRcs.MA2025-01中稿 · SIGIR-AP 2025被引 6

让AI像人一样规划检索,更准更快生成多模态内容。

CogPlanner: Unveiling the Potential of Agentic Multimodal Retrieval Augmented Generation with Planning

  • 模仿人类认知,分步规划检索与提问策略。
  • 在多个数据集上准确率提升显著,计算开销几乎不变。
  • 适合需要高效多模态生成的智能系统开发者。

多模态检索增强生成(MRAG)系统在提升多模态大语言模型(MLLM)生成能力方面展现出潜力。然而,现有框架大多采用固定的一次性检索策略,难以应对信息获取与查询重构等真实场景挑战。本文提出多模态检索增强生成规划(MRAG Planning)任务,旨在有效寻获信息并整合输出,同时最小化计算开销。我们提出CogPlanner,一个受人类认知启发的代理式即插即用框架,通过迭代决策查询重构与检索策略,生成准确且上下文相关的回答。该框架支持并行与串行建模范式。此外,我们构建了CogBench基准,用于严格评估MRAG Planning任务,并促进CogPlanner与轻量级MLLM(如Qwen2-VL-7B-Cog)的高效集成。实验表明,CogPlanner显著优于现有MRAG基线,在准确率和效率上均有提升,且额外计算成本极低。

原文摘要 · Abstract (English)

Multimodal Retrieval Augmented Generation (MRAG) systems have shown promise in enhancing the generation capabilities of multimodal large language models (MLLMs). However, existing MRAG frameworks primarily adhere to rigid, single-step retrieval strategies that fail to address real-world challenges of information acquisition and query reformulation. In this work, we introduce the task of Multimodal Retrieval Augmented Generation Planning (MRAG Planning) that aims at effective information seeking and integration while minimizing computational overhead. Specifically, we propose CogPlanner, an agentic plug-and-play framework inspired by human cognitive processes, which iteratively determines query reformulation and retrieval strategies to generate accurate and contextually relevant responses. CogPlanner supports parallel and sequential modeling paradigms. Furthermore, we introduce CogBench, a new benchmark designed to rigorously evaluate the MRAG Planning task and facilitate lightweight CogPlanner integration with resource-efficient MLLMs, such as Qwen2-VL-7B-Cog. Experimental results demonstrate that CogPlanner significantly outperforms existing MRAG baselines, offering improvements in both accuracy and efficiency with minimal additional computational costs.

多模态生成智能规划检索增强轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。