arXiv:2504.15800cs.IR2025-04被引 29

构建金融领域真实检索问答数据集,推动可信生成研究

FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation

  • 专家标注5703组查询-证据-答案三元组,贴近真实金融搜索行为
  • 挑战模型从海量文档中精准检索,而非依赖预设上下文
  • 适合研究金融领域检索增强生成的学者与工程师

在快速变化的金融领域,准确及时的信息对应对市场动态至关重要。正确检索信息是金融问答的关键,因为许多语言模型在此领域难以保证事实准确性。本文提出FinDER,一个专为金融领域检索增强生成(RAG)设计的专家生成数据集。不同于现有提供预设上下文且问题清晰的QA数据集,FinDER聚焦于由领域专家标注的搜索相关证据,包含5703个源自真实金融咨询的查询-证据-答案三元组。这些查询常含缩写、首字母词和简洁表达,反映专业人士真实搜索行为中的简略性和模糊性。通过要求模型从大规模语料库中检索相关信息,而非依赖预设上下文,FinDER为评估RAG系统提供了更真实的基准。我们进一步对多种先进检索模型和大语言模型进行了全面评估,揭示了在真实场景下存在的挑战,以推动金融领域可信、精准RAG技术的未来发展。

原文摘要 · Abstract (English)

In the fast-paced financial domain, accurate and up-to-date information is critical to addressing ever-evolving market conditions. Retrieving this information correctly is essential in financial Question-Answering (QA), since many language models struggle with factual accuracy in this domain. We present FinDER, an expert-generated dataset tailored for Retrieval-Augmented Generation (RAG) in finance. Unlike existing QA datasets that provide predefined contexts and rely on relatively clear and straightforward queries, FinDER focuses on annotating search-relevant evidence by domain experts, offering 5,703 query-evidence-answer triplets derived from real-world financial inquiries. These queries frequently include abbreviations, acronyms, and concise expressions, capturing the brevity and ambiguity common in the realistic search behavior of professionals. By challenging models to retrieve relevant information from large corpora rather than relying on readily determined contexts, FinDER offers a more realistic benchmark for evaluating RAG systems. We further present a comprehensive evaluation of multiple state-of-the-art retrieval models and Large Language Models, showcasing challenges derived from a realistic benchmark to drive future research on truthful and precise RAG in the financial domain.

金融AI检索增强数据集问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。