构建日文财报评测集,检验大模型在金融复杂任务中的真实能力
EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements
- 基于十年日本公司年报构建多任务评测集
- 顶尖大模型在欺诈检测等任务上仅略优于逻辑回归
- 揭示现有评测需更贴近真实金融工作场景
大型语言模型在数学、编程等领域已超越人类表现,其进步得益于基准数据集的建设。然而,金融领域因专业门槛高,相关评测数据集仍相对匮乏。本文提出EDINET-Bench,一个开源的日文金融评测基准,用于评估大模型在会计欺诈检测、盈利预测和行业分类等复杂任务上的表现。该数据集源自十年间日本企业提交的年度报告,要求模型处理整份报告,整合多个表格与文本段落的信息,具备专家级推理能力,对人类专业人士亦具挑战性。实验表明,即使是最先进的大模型在此领域表现依然有限,在二分类任务中仅略优于逻辑回归。这说明直接提供报告给模型的简单设定不足以发挥其潜力,亟需更贴近金融从业者实际工作环境的评测框架,如真实模拟和任务专用推理支持。数据集与代码已公开,以推动后续研究。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have made remarkable progress, surpassing human performance on several benchmarks in domains such as mathematics and coding. A key driver of this progress has been the development of benchmark datasets. In contrast, the financial domain poses higher entry barriers due to its demand for specialized expertise, and benchmarks remain relatively scarce compared to those in mathematics or coding. We introduce EDINET-Bench, an open-source Japanese financial benchmark designed to evaluate LLMs on challenging tasks such as accounting fraud detection, earnings forecasting, and industry classification. EDINET-Bench is constructed from ten years of annual reports filed by Japanese companies. These tasks require models to process entire annual reports and integrate information across multiple tables and textual sections, demanding expert-level reasoning that is challenging even for human professionals. Our experiments show that even state-of-the-art LLMs struggle in this domain, performing only marginally better than logistic regression in binary classification tasks such as fraud detection and earnings forecasting. Our results show that simply providing reports to LLMs in a straightforward setting is not enough. This highlights the need for benchmark frameworks that better reflect the environments in which financial professionals operate, with richer scaffolding such as realistic simulations and task-specific reasoning support to enable more effective problem solving. We make our dataset and code publicly available to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。