arXiv:2411.16365cs.CL2024-11被引 10

构建多模态检索增强生成的基准,提升模型对网络多模态内容的理解与生成能力

Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines

  • 设计数据筛选流程,构建高质量多模态检索生成数据集
  • 7B-8B微调模型性能超越GPT-4o,接近o3-mini水平
  • 提出基于基础模型的多模态评估指标,支持跨领域分析

我们系统研究了多模态检索增强多模态生成(M²RAG)这一新任务,使大模型能够处理多模态网络内容并生成多模态响应,具有更高的信息密度和可读性。尽管潜力巨大,该任务仍缺乏深入研究和高质量数据资源。为此,我们通过严格的清洗流程建立全面基准,采用基于基础模型的文本-模态与多模态评估指标进行评测,并提出有效策略以帮助大模型完成任务,通过自定义指标筛选高质量样本构建训练集。大量实验验证了评估指标的可靠性,揭示了不同策略下的模型表现格局,表明微调后的7B-8B模型优于GPT-4o,接近OpenAI o3-mini。进一步细粒度分析覆盖多个领域,证实了数据构建流程的有效性。所有资源(代码、数据集、模型权重)将公开发布。

原文摘要 · Abstract (English)

We present a systematic investigation of Multi-modal Retrieval Augmented Multi-modal Generation (M$^2$RAG), a novel task that enables foundation models to process multi-modal web content and generate multi-modal responses, which exhibits better information density and readability. Despite its potential impact, M$^2$RAG remains understudied, lacking comprehensive analysis and high-quality data resources. To address this gap, we establish a comprehensive benchmark through a rigorous data curation pipeline, and employ text-modal metrics and multi-modal metrics based on foundation models for evaluation. We further propose several strategies for foundation models to process M$^2$RAG task effectively and construct a training set by filtering high-quality samples using our designed metrics. Our extensive experiments demonstrate the reliability of our proposed metrics, a landscape of model performance within our designed strategies, and show that our fine-tuned 7B-8B models outperform the GPT-4o model and approach the state-of-the-art OpenAI o3-mini. Additionally, we perform fine-grained analyses across diverse domains and validate the effectiveness of our designs in data curation pipeline. All resources, including codes, datasets, and model weights, will be publicly released.

多模态生成检索增强基准测试大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。