arXiv:2602.17106cs.AI2026-02被引 1

构建人机协作框架,提升企业可持续性评级的可信度与可比性。

Toward Trustworthy Evaluation of Sustainability Rating Methodologies: A Human-AI Collaborative Framework for Benchmark Dataset Construction

  • 提出人机协同框架STRIDE,用大模型生成标准化评级数据集。
  • 通过SR-Delta分析差异,发现评级不一致的根源并提出优化方向。
  • 为评估可持续性评级方法提供可扩展、可比较的基准数据支持。

可持续性或ESG评级机构利用企业披露信息和外部数据,对企业的环境、社会与治理表现进行评分。然而,不同机构对同一公司的评级差异巨大,影响其可比性、可信度及决策参考价值。为实现评级结果的统一,本文提出采用通用的人机协作框架,构建可信的基准数据集以评估可持续性评级方法。该框架包含两部分:STRIDE(可持续性可信评级与完整性数据方程)提供原则性标准与评分体系,指导基于大语言模型(LLMs)的企业级基准数据集构建;SR-Delta则是一种差异分析流程框架,可揭示潜在调整建议。该框架实现了可持续性评级方法的可扩展、可比较评估。本文呼吁更广泛的AI研究社区采用智能化方法,推动可持续性评级方法的强化与演进,以支持紧迫的可持续发展目标。

原文摘要 · Abstract (English)

Sustainability or ESG rating agencies use company disclosures and external data to produce scores or ratings that assess the environmental, social, and governance performance of a company. However, sustainability ratings across agencies for a single company vary widely, limiting their comparability, credibility, and relevance to decision-making. To harmonize the rating results, we propose adopting a universal human-AI collaboration framework to generate trustworthy benchmark datasets for evaluating sustainability rating methodologies. The framework comprises two complementary parts: STRIDE (Sustainability Trust Rating & Integrity Data Equation) provides principled criteria and a scoring system that guide the construction of firm-level benchmark datasets using large language models (LLMs), and SR-Delta, a discrepancy-analysis procedural framework that surfaces insights for potential adjustments. The framework enables scalable and comparable assessment of sustainability rating methodologies. We call on the broader AI community to adopt AI-powered approaches to strengthen and advance sustainability rating methodologies that support and enforce urgent sustainability agendas.

可持续性评级人机协作大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。