用半合成方法构建金融领域评测基准,解决高质量数据稀缺问题。
FinForge: Semi-Synthetic Financial Benchmark Generation
- 结合专家标注与大模型生成,构建金融领域问答数据
- 产出5000+经人工验证的题目,覆盖11个金融子领域
- 可评估模型金融推理能力,适合研究金融AI的学者使用
在金融等高风险专业领域评估语言模型仍面临挑战,主要因缺乏公开、高质量且领域专精的数据集。现有通用基准覆盖面广但深度不足,难以衡量模型在真实金融推理中所需的概念理解与量化能力。为此,我们提出FinForge,一种可扩展的半合成基准生成管道,融合专家引导的数据筛选与受控的语言模型合成。该方法通过权威金融资料的文档采集与程序化构建,结合Gemini 2.5 Flash进行结构化问题生成与验证。为验证其有效性,我们构建了包含超过5,000个经人工验证的问答对的FinForge-5k基准,数据源自10万份经核实的文档,总计143M tokens,覆盖11个金融子领域。在FinForge-5k上对主流开源与闭源模型的评估显示,领先模型准确率接近80%,显著揭示了当前模型在金融推理中的能力差异。该框架有助于诊断模型缺陷并推动金融领域智能系统的改进。所有代码与数据均可在https://github.com/gtfintechlab/FinForge获取。
原文摘要 · Abstract (English)
Evaluating Language Models (LMs) in specialized, high-stakes domains such as finance remains a significant challenge due to the scarcity of open, high-quality, and domain-specific datasets. Existing general-purpose benchmarks provide broad coverage but lack the depth and domain fidelity needed to assess LMs' capabilities for real-world financial reasoning, which requires both conceptual understanding and quantitative rigor. To address this gap, we introduce FinForge, a scalable, semi-synthetic pipeline for constructing finance-specific evaluation benchmarks through a hybrid of expert-guided data curation and controlled LM-based synthesis. FinForge combines manual and programmatic corpus construction from authoritative financial sources with structured question generation and validation using Gemini 2.5 Flash. To demonstrate the pipeline's efficacy, we produce FinForge-5k, a snapshot benchmark comprising over 5,000 human-validated question-answer pairs across 11 finance subdomains, derived from a curated corpus of 100,000 verified documents totaling 143M tokens. Evaluation of state-of-the-art open-source and closed-source models on FinForge-5k reveals significant differences in financial reasoning, with leading models achieving accuracy levels near 80%. These findings underscore the framework's utility for diagnosing current model limitations and guiding future improvements in financial domain competence. All code and data are available at https://github.com/gtfintechlab/FinForge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。