arXiv:2506.01646cs.CLcs.AI2025-06EMNLP被引 14

首个系统评估大模型在可持续发展领域的问答能力基准。

ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge

  • 构建1136道经专家验证的多选题,链接原始文献支持透明评估。
  • 小模型用RAG后准确率从63.8%提升至80.5%,显著改善知识缺失。
  • 适合关注绿色金融、ESG合规与可信AI研究者使用。

我们提出ESGenius,一个全面评估和提升大语言模型(LLMs)在环境、社会与治理(ESG)及可持续发展领域问答能力的基准。该基准包含两个核心部分:(i) ESGenius-QA,由大模型生成并经领域专家严格验证的1,136道多选题,覆盖广泛ESG议题与可持续发展主题,每道题均关联原始文本来源,支持可解释评估与检索增强生成(RAG)方法;(ii) ESGenius-Corpus,从7个权威机构精选的231份基础框架、标准、报告与建议文档。为全面评估模型能力与适应潜力,采用零样本与RAG两阶段评估协议。对50种不同规模的模型(0.5B至671B)进行实验表明,先进模型在零样本下仅达55–70%准确率,暴露其在跨学科专业领域存在显著知识空白。而采用RAG的模型表现大幅提升,如DeepSeek-R1-Distill-Qwen-14B从63.82%升至80.46%。结果表明,基于权威来源的外部知识增强对提升模型在可持续发展理解至关重要。据我们所知,ESGenius是首个专为严谨评估大模型在ESG与可持续发展知识方面能力而设计的综合性基准,为推动该关键领域可信AI发展提供重要工具。

原文摘要 · Abstract (English)

We introduce ESGenius, a comprehensive benchmark for evaluating and enhancing the proficiency of Large Language Models (LLMs) in Environmental, Social, and Governance (ESG) and sustainability-focused question answering. ESGenius comprises two key components: (i) ESGenius-QA, a collection of 1,136 Multiple-Choice Questions (MCQs) generated by LLMs and rigorously validated by domain experts, covering a broad range of ESG pillars and sustainability topics. Each question is systematically linked to its corresponding source text, enabling transparent evaluation and supporting Retrieval-Augmented Generation (RAG) methods; and (ii) ESGenius-Corpus, a meticulously curated repository of 231 foundational frameworks, standards, reports, and recommendation documents from 7 authoritative sources. Moreover, to fully assess the capabilities and adaptation potential of LLMs, we implement a rigorous two-stage evaluation protocol -- Zero-Shot and RAG. Extensive experiments across 50 LLMs (0.5B to 671B) demonstrate that state-of-the-art models achieve only moderate performance in zero-shot settings, with accuracies around 55--70%, highlighting a significant knowledge gap for LLMs in this specialized, interdisciplinary domain. However, models employing RAG demonstrate significant performance improvements, particularly for smaller models. For example, DeepSeek-R1-Distill-Qwen-14B improves from 63.82% (zero-shot) to 80.46% with RAG. These results demonstrate the necessity of grounding responses in authoritative sources for enhanced ESG understanding. To the best of our knowledge, ESGenius is the first comprehensive QA benchmark designed to rigorously evaluate LLMs on ESG and sustainability knowledge, providing a critical tool to advance trustworthy AI in this vital domain.

ESG大模型评测可持续发展RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。