arXiv:2506.02037cs.CLcs.AI2025-06被引 2

构建首个面向金融领域在线RAG系统的评测基准,解决数据实时性与保密性难题。

FinS-Pilot: A Benchmark for Online Financial RAG System

  • 基于真实金融助手对话,融合实时API与文本数据
  • 覆盖关键金融场景,支持静态知识与动态市场信息评估
  • 适配中文大模型,助力金融AI系统选型与优化

大型语言模型在多个专业领域展现出卓越能力,其性能通常通过标准化基准进行评估。在金融领域,对专业准确性与实时数据处理的严苛要求常需依赖检索增强生成(RAG)技术。然而,金融RAG基准的发展受限于数据保密问题及动态数据整合缺失。为此,我们提出FinS-Pilot,一个面向在线金融应用RAG系统的新型评测基准。该基准基于真实金融助手交互构建,融合实时API数据与文本数据,通过意图分类框架覆盖关键金融领域。其可全面评估金融助手在处理静态知识与时效性市场信息方面的能力。通过对多款中国主流大模型的系统性实验,验证了FinS-Pilot在识别适用于金融场景模型方面的有效性,弥补了金融领域专用评估工具的空白。本研究贡献了一个实用评估框架与精炼数据集,推动金融NLP系统研究发展。代码与数据已公开于GitHub。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities across various professional domains, with their performance typically evaluated through standardized benchmarks. In the financial field, the stringent demands for professional accuracy and real-time data processing often necessitate the use of retrieval-augmented generation (RAG) techniques. However, the development of financial RAG benchmarks has been constrained by data confidentiality issues and the lack of dynamic data integration. To address this issue, we introduce FinS-Pilot, a novel benchmark for evaluating RAG systems in online financial applications. Constructed from real-world financial assistant interactions, our benchmark incorporates both real-time API data and text data, organized through an intent classification framework covering critical financial domains. The benchmark enables comprehensive evaluation of financial assistants' capabilities in handling both static knowledge and time-sensitive market information.Through systematic experiments with multiple Chinese leading LLMs, we demonstrate FinS-Pilot's effectiveness in identifying models suitable for financial applications while addressing the current gap in specialized evaluation tools for the financial domain. Our work contributes both a practical evaluation framework and a curated dataset to advance research in financial NLP systems. The code and dataset are accessible on GitHub.

金融AIRAG评测大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。