用大模型评估企业碳披露,发现能识别优秀表现者且抗伪饰。
Judging It, Washing It: Scoring and Greenwashing Corporate Climate Disclosures using Large Language Models
- 用大模型评分企业减排目标与进展,分数值与排名有效。
- 大模型可生成符合要求的伪环保文案,存在绿色洗白风险。
- 配对比较评分法更抗伪造,适合用于监督企业披露质量。
我们研究了大型语言模型(LLMs)在评估和绿色洗白企业碳排放披露中的应用。首先,探讨了使用大模型作为评审员(LLM-as-a-Judge, LLMJ)方法对提交的减排目标与进展报告进行评分的有效性。其次,探究了在准确性与长度约束下,大模型生成绿色洗白回应的行为特征。最后,测试了LLMJ方法在面对经大模型绿色洗白后的回应时的鲁棒性。结果表明,两种LLMJ评分系统——数值评分与配对比较——均能有效区分表现优异的企业,其中配对比较系统对大模型生成的绿色洗白内容表现出更强的鲁棒性。
原文摘要 · Abstract (English)
We study the use of large language models (LLMs) to both evaluate and greenwash corporate climate disclosures. First, we investigate the use of the LLM-as-a-Judge (LLMJ) methodology for scoring company-submitted reports on emissions reduction targets and progress. Second, we probe the behavior of an LLM when it is prompted to greenwash a response subject to accuracy and length constraints. Finally, we test the robustness of the LLMJ methodology against responses that may be greenwashed using an LLM. We find that two LLMJ scoring systems, numerical rating and pairwise comparison, are effective in distinguishing high-performing companies from others, with the pairwise comparison system showing greater robustness against LLM-greenwashed responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。