arXiv:2508.08816cs.AI2025-08被引 1

提出E-Agent框架,让多模态检索生成更智能高效。

Efficient Agent: Optimizing Planning Capability for Multimodal Retrieval Augmented Generation

  • 用动态规划器根据上下文调度多模态工具
  • 比顶尖方法准确率高13%,冗余搜索减少37%
  • 适合需要实时分析新闻、热点的开发者

多模态检索增强生成(mRAG)已成为解决多模态大模型在新闻分析、热点追踪等真实场景中时效性瓶颈的有前景方案。然而现有方法常受限于僵化检索策略和视觉信息利用不足。为此,我们提出E-Agent,一个包含两大创新的智能体框架:一个基于上下文推理动态编排多模态工具的mRAG规划器,以及一个采用工具感知执行顺序的任务执行器,以实现优化的mRAG工作流。E-Agent采用一次性mRAG规划策略,在提升信息检索效率的同时减少冗余工具调用。为系统评估mRAG系统的规划能力,我们引入了真实世界mRAG规划基准(RemPlan)。该基准包含依赖与不依赖检索的问题类型,对每例进行必要检索工具的系统标注。其显式的mRAG规划标注与多样化问题设计增强了实际应用相关性,可模拟需动态决策的真实场景。在RemPlan及三个已有基准上的实验表明,E-Agent表现卓越:相比最先进mRAG方法,准确率提升13%,冗余搜索减少37%。

原文摘要 · Abstract (English)

Multimodal Retrieval-Augmented Generation (mRAG) has emerged as a promising solution to address the temporal limitations of Multimodal Large Language Models (MLLMs) in real-world scenarios like news analysis and trending topics. However, existing approaches often suffer from rigid retrieval strategies and under-utilization of visual information. To bridge this gap, we propose E-Agent, an agent framework featuring two key innovations: a mRAG planner trained to dynamically orchestrate multimodal tools based on contextual reasoning, and a task executor employing tool-aware execution sequencing to implement optimized mRAG workflows. E-Agent adopts a one-time mRAG planning strategy that enables efficient information retrieval while minimizing redundant tool invocations. To rigorously assess the planning capabilities of mRAG systems, we introduce the Real-World mRAG Planning (RemPlan) benchmark. This novel benchmark contains both retrieval-dependent and retrieval-independent question types, systematically annotated with essential retrieval tools required for each instance. The benchmark's explicit mRAG planning annotations and diverse question design enhance its practical relevance by simulating real-world scenarios requiring dynamic mRAG decisions. Experiments across RemPlan and three established benchmarks demonstrate E-Agent's superiority: 13% accuracy gain over state-of-the-art mRAG methods while reducing redundant searches by 37%.

多模态检索增强智能体效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。