arXiv:2505.17471cs.CL2025-05EMNLP被引 20

金融领域首个支持图文引用的多模态检索增强生成基准

FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain

  • 构建包含中英双语文档的多模态金融数据集
  • 提出RGenCite模型实现图文联合生成与引用
  • 设计自动评估方法验证多模态大模型的引用能力

检索增强生成(RAG)在金融领域至关重要,支撑实时市场分析、趋势预测和利率计算等应用。然而现有研究主要聚焦文本数据,忽视了金融文档中的丰富视觉内容,导致关键分析洞察丢失。为此,我们提出FinRAGBench-V,一个面向金融领域的综合性多模态RAG基准,有效融合多源数据并提供可视化引用以确保可追溯性。该基准包含60,780页中文与51,219页英文的双语检索语料库,以及涵盖七类问题、多种数据类型的高质量人工标注问答数据集。我们还引入RGenCite基线模型,实现生成与视觉引用的无缝集成,并提出一种自动引用评估方法,系统评估多模态大语言模型的引用能力。在RGenCite上的广泛实验揭示了FinRAGBench-V的挑战性,为金融领域多模态RAG系统的发展提供了宝贵洞见。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) plays a vital role in the financial domain, powering applications such as real-time market analysis, trend forecasting, and interest rate computation. However, most existing RAG research in finance focuses predominantly on textual data, overlooking the rich visual content in financial documents, resulting in the loss of key analytical insights. To bridge this gap, we present FinRAGBench-V, a comprehensive visual RAG benchmark tailored for finance which effectively integrates multimodal data and provides visual citation to ensure traceability. It includes a bilingual retrieval corpus with 60,780 Chinese and 51,219 English pages, along with a high-quality, human-annotated question-answering (QA) dataset spanning heterogeneous data types and seven question categories. Moreover, we introduce RGenCite, an RAG baseline that seamlessly integrates visual citation with generation. Furthermore, we propose an automatic citation evaluation method to systematically assess the visual citation capabilities of Multimodal Large Language Models (MLLMs). Extensive experiments on RGenCite underscore the challenging nature of FinRAGBench-V, providing valuable insights for the development of multimodal RAG systems in finance.

多模态RAG金融AI视觉引用基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。