构建开放基准,用大模型问答企业碳披露信息。
Climate Finance Bench
- 基于33份跨11个行业的可持续报告,构建330个专家验证的问答对。
- 发现检索模块定位答案段落的能力是性能瓶颈。
- 倡导AI气候应用中透明碳披露,推荐量化等技术。
Climate Finance Bench 提出一个开放基准,用于评估大语言模型在企业气候披露文本上的问答能力。我们从全部11个GICS行业精选33份近期英文可持续发展报告,并标注了330个经专家验证的问答对,涵盖纯信息提取、数值推理与逻辑推理。基于该数据集,我们对比了多种RAG(检索增强生成)方法,发现检索器准确找到含答案段落的能力是性能关键瓶颈。研究进一步呼吁在气候相关AI应用中推行透明碳披露,强调权重量化等技术的优势。
原文摘要 · Abstract (English)
Climate Finance Bench introduces an open benchmark that targets question-answering over corporate climate disclosures using Large Language Models. We curate 33 recent sustainability reports in English drawn from companies across all 11 GICS sectors and annotate 330 expert-validated question-answer pairs that span pure extraction, numerical reasoning, and logical reasoning. Building on this dataset, we propose a comparison of RAG (retrieval-augmented generation) approaches. We show that the retriever's ability to locate passages that actually contain the answer is the chief performance bottleneck. We further argue for transparent carbon reporting in AI-for-climate applications, highlighting advantages of techniques such as Weight Quantization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。