测试大模型在真实金融分析任务中的表现,发现现有模型准确率不足40%。
FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models
- 构建真实金融场景的多步数值分析任务评估体系
- 模型在模拟投行等机构任务中失败率约60%
- 适合研究金融AI、大模型量化推理能力的学者使用
FinanceQA 是一个评估大语言模型在复杂数值金融分析任务中表现的测试套件,模拟对冲基金、私募股权公司、投资银行等金融机构的真实工作场景。尽管近期有进展,当前模型仍难以满足金融机构对准确性的严格要求,约60%的现实任务无法正确完成。主要挑战包括手动展开财务指标、遵循标准会计与企业估值规范,以及在信息不全条件下进行多步骤分析并生成假设。该表现差距凸显了现有模型能力与专业金融分析需求之间的脱节,也表明需要更高质量训练数据支持此类任务。研究通过OpenAI微调接口进行了初步实验。FinanceQA 已在 Hugging Face 公开发布。
原文摘要 · Abstract (English)
FinanceQA is a testing suite that evaluates LLMs' performance on complex numerical financial analysis tasks that mirror real-world investment work. Despite recent advances, current LLMs fail to meet the strict accuracy requirements of financial institutions, with models failing approximately 60% of realistic tasks that mimic on-the-job analyses at hedge funds, private equity firms, investment banks, and other financial institutions. The primary challenges include hand-spreading metrics, adhering to standard accounting and corporate valuation conventions, and performing analysis under incomplete information - particularly in multi-step tasks requiring assumption generation. This performance gap highlights the disconnect between existing LLM capabilities and the demands of professional financial analysis that are inadequately tested by current testing architectures. Results show that higher-quality training data is needed to support such tasks, which we experiment with using OpenAI's fine-tuning API. FinanceQA is publicly released at [this https URL](https://huggingface.co/datasets/AfterQuery/FinanceQA).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。