用多智能体生成复杂多模态问答数据,评估企业级检索增强模型
MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation
- 构建多智能体协作系统,模拟专家推理流程生成高质量数据
- 生成数据平均需2.3次以上跳转推理,事实准确性显著提升
- 适合研究检索增强生成、多模态推理及领域专用评估的团队
检索增强生成(RAG)在多模态、高风险的企业应用中快速发展,但领域专用评估基准却未能同步跟进。现有数据集多基于通用语料或纯文本检索,难以捕捉专业文档中信息高度多模态且需整合离散证据的复杂性。为此,我们提出MiRAGE——一个用于RAG系统评估的多智能体框架,通过协作式智能体群生成经验证的、领域特定的、多模态且多跳推理的问答数据集。MiRAGE协调多个专用智能体:递归上下文优化循环以聚合分散证据,对抗性验证智能体确保事实一致性,以及识别专家角色与相关领域的智能体,以模仿专家认知流程。在法规、金融、量化生物学和新闻四个不同领域进行的实证评估表明,MiRAGE生成的数据平均推理跳数超过2.3次,且事实忠实度显著更高。消融实验显示,若图像有文本描述,可用大语言模型驱动该框架;但视觉定位仍是前沿挑战。通过自动化创建反映专有语料潜在主题结构的黄金标准评估数据集,MiRAGE为严格评测下一代信息检索系统提供了必要基础设施。
原文摘要 · Abstract (English)
The rapid evolution of Retrieval-Augmented Generation (RAG) toward multimodal, high-stakes enterprise applications has outpaced the development of domain specific evaluation benchmarks. Existing datasets often rely on general-domain corpora or purely textual retrieval, failing to capture the complexity of specialized technical documents where information is inextricably multimodal and reasoning requires synthesizing disjoint evidence. We address this gap by introducing MiRAGE, a Multiagent framework for RAG systems Evaluation, that leverages a collaborative swarm of specialized agents to generate verified, domain-specific, multimodal, and multi-hop Question-Answer datasets. MiRAGE orchestrates a swarm of specialized agents: a recursive context optimization loop to aggregate scattered evidence, an adversarial verifier agent to guarantee factual grounding, and an agent to recognize the expert persona and the relevant domain to mimic expert cognitive workflows. Extensive empirical evaluation across four distinct domains (regulations, finance, quantitative biology, and journalism) demonstrates that MiRAGE generates datasets with significantly higher reasoning complexity (>2.3 average hops) and factual faithfulness. Our ablation studies point that MiRAGE can be powered by LLMs if textual descriptions of the images are available. Visual grounding still remains a frontier. By automating the creation of gold standard evaluation datasets that reflect the latent thematic structure of proprietary corpora, MiRAGE provides the necessary infrastructure to rigorously benchmark the next generation information retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。