构建金融分析评测基准,评估大模型在真实财报中的表现
SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities
- 设计565道专家撰写的问题,覆盖四大财务分析任务
- 用多模型评判机制验证,结果与人工评估高度一致
- 适合研究金融AI的学者和开发者使用
我们提出SECQUE,一个全面的大型语言模型(LLM)金融分析能力评测基准。SECQUE包含565道专家撰写的题目,涵盖四大核心领域:对比分析、比率计算、风险评估和财务洞察生成,均基于美国证券交易委员会(SEC)文件。为评估模型表现,我们开发了SECQUE-Judge,一种基于多个LLM的评判机制,其结果与人类评价具有强一致性。此外,我们对多种模型在该基准上的表现进行了深入分析。通过公开发布SECQUE,我们希望推动金融人工智能领域的进一步研究与进步。
原文摘要 · Abstract (English)
We introduce SECQUE, a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks. SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories: comparison analysis, ratio calculation, risk assessment, and financial insight generation. To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human evaluations. Additionally, we provide an extensive analysis of various models' performance on our benchmark. By making SECQUE publicly available, we aim to facilitate further research and advancements in financial AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。