arXiv:2510.24402cs.IRcs.AI2025-10被引 7

用元数据增强金融文档问答,提升检索准确率。

Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering

  • 利用大模型生成元数据,构建上下文丰富的文本块。
  • 结合元数据嵌入与重排序,使金融问答准确率显著提升。
  • 适合金融分析、合规审查等需要精准理解财报的场景。

在长篇结构化财务文件中,相关证据稀疏且跨章节引用,传统检索增强生成(RAG)表现不佳。本文系统研究了先进的元数据驱动RAG技术,提出并评估一种多阶段RAG架构,该架构依赖大模型生成的元数据。设计了复杂的索引流程,创建语义丰富文档块,并在FinanceBench数据集上对比多种优化方法,包括预检索过滤、后检索重排序和增强嵌入。结果表明,虽然强重排序器对精度至关重要,但最大性能提升来自将元数据与文本直接融合形成“上下文块”。最优架构结合大模型预处理与上下文嵌入,实现卓越效果。此外,我们提出一种自研元数据重排序器,为商业方案提供高性价比替代,平衡性能与效率。本研究为构建稳健的金融文档分析RAG系统提供了可复现蓝图。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) struggles on long, structured financial filings where relevant evidence is sparse and cross-referenced. This paper presents a systematic investigation of advanced metadata-driven Retrieval-Augmented Generation (RAG) techniques, proposing and evaluating a novel, multi-stage RAG architecture that leverages LLM-generated metadata. We introduce a sophisticated indexing pipeline to create contextually rich document chunks and benchmark a spectrum of enhancements, including pre-retrieval filtering, post-retrieval reranking, and enriched embeddings, benchmarked on the FinanceBench dataset. Our results reveal that while a powerful reranker is essential for precision, the most significant performance gains come from embedding chunk metadata directly with text ("contextual chunks"). Our proposed optimal architecture combines LLM-driven pre-retrieval optimizations with these contextual embeddings to achieve superior performance. Additionally, we present a custom metadata reranker that offers a compelling, cost-effective alternative to commercial solutions, highlighting a practical trade-off between peak performance and operational efficiency. This study provides a blueprint for building robust, metadata-aware RAG systems for financial document analysis.

金融问答元数据RAG大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。