构建可解释的可持续发展问答基准,评估模型推理能力
ESGBench: A Benchmark for Explainable ESG Question Answering in Corporate Sustainability Reports
- 基于企业可持续报告设计多主题可解释问答任务
- 揭示主流大模型在事实一致性与领域对齐上的缺陷
- 适合关注绿色金融与可信AI的研究者使用
我们提出ESGBench,一个用于评估基于企业可持续报告的可解释ESG问答系统的基准数据集与评估框架。该基准包含跨多个ESG主题的领域相关问题,配有经过人工校准的答案及支持证据,以实现对模型推理过程的细粒度评估。我们分析了当前主流大模型在ESGBench上的表现,揭示了其在事实一致性、可追溯性和领域适配性方面的关键挑战。ESGBench旨在推动透明且可问责的可持续发展导向AI系统研究。
原文摘要 · Abstract (English)
We present ESGBench, a benchmark dataset and evaluation framework designed to assess explainable ESG question answering systems using corporate sustainability reports. The benchmark consists of domain-grounded questions across multiple ESG themes, paired with human-curated answers and supporting evidence to enable fine-grained evaluation of model reasoning. We analyze the performance of state-of-the-art LLMs on ESGBench, highlighting key challenges in factual consistency, traceability, and domain alignment. ESGBench aims to accelerate research in transparent and accountable ESG-focused AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。