用大模型自动分析企业可持续报告,发现定量预测仍不靠谱。
Automated Analysis of Sustainability Reports: Using Large Language Models for the Extraction and Prediction of EU Taxonomy-Compliant KPIs
- 构建190份企业报告的结构化数据集,用于评估大模型合规分析能力。
- 大模型能勉强识别经济活动,但零样本预测财务指标几乎失败。
- 简洁元数据比完整报告表现更好,适合做人工辅助工具。
企业遵守欧盟可持续性分类标准(EU Taxonomy)的流程依赖人工,耗时耗力。尽管大语言模型(LLMs)提供了自动化潜力,但缺乏公开基准数据集限制了研究进展。为此,我们构建了一个包含190份企业报告的新结构化数据集,涵盖真实经济活动和定量关键绩效指标(KPIs)。基于此数据集,我们首次系统评估了大模型在核心合规工作流中的表现。结果表明,定性任务与定量任务之间存在明显差距:大模型在识别经济活动方面表现中等,多步代理框架略有提升精度;但在零样本设置下,对财务类KPI的预测全面失败。我们还发现悖论现象——简洁的元数据往往优于完整非结构化报告,且模型置信度评分严重失准。结论是,当前大模型尚不足以实现全自动化,但可作为人类专家的强力辅助工具。本研究提供的数据集为未来研究提供了公开基准。
原文摘要 · Abstract (English)
The manual, resource-intensive process of complying with the EU Taxonomy presents a significant challenge for companies. While Large Language Models (LLMs) offer a path to automation, research is hindered by a lack of public benchmark datasets. To address this gap, we introduce a novel, structured dataset from 190 corporate reports, containing ground-truth economic activities and quantitative Key Performance Indicators (KPIs). We use this dataset to conduct the first systematic evaluation of LLMs on the core compliance workflow. Our results reveal a clear performance gap between qualitative and quantitative tasks. LLMs show moderate success in the qualitative task of identifying economic activities, with a multi-step agentic framework modestly enhancing precision. Conversely, the models comprehensively fail at the quantitative task of predicting financial KPIs in a zero-shot setting. We also discover a paradox, where concise metadata often yields superior performance to full, unstructured reports, and find that model confidence scores are poorly calibrated. We conclude that while LLMs are not ready for full automation, they can serve as powerful assistive tools for human experts. Our dataset provides a public benchmark for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。