arXiv:2601.08676cs.AI2026-01

构建智能代理与评测体系,提升企业可持续性分析的准确性与专业性

Advancing ESG Intelligence: An Expert-level Agent and Comprehensive Benchmark for Sustainable Finance

  • 采用分层多智能体架构,结合检索增强与领域工具完成复杂分析任务
  • 在原子问答任务中准确率达84.15%,报告生成融合图表与可验证引用
  • 适用于金融风控、ESG评级等高要求场景的专业智能分析

环境、社会与治理(ESG)标准是评估企业可持续性与道德表现的关键。然而,专业级ESG分析受限于非结构化数据分散问题,现有大语言模型难以处理复杂的多步审计流程。为此,我们提出ESGAgent——一种由专用工具集支持的分层多智能体系统,包含检索增强、网络搜索及领域特定功能,可生成深度ESG分析。同时,我们基于310份企业可持续性报告构建了三级综合评测体系,涵盖从基础常识问答到整合性深度分析的能力评估。实证表明,ESGAgent在原子问答任务中平均准确率达84.15%,显著优于主流闭源LLM;在专业报告生成中能有效集成丰富图表与可验证引用。结果证实该评测体系具备诊断价值,可作为高风险垂直领域中通用与进阶智能体能力评估的重要基准。

原文摘要 · Abstract (English)

Environmental, social, and governance (ESG) criteria are essential for evaluating corporate sustainability and ethical performance. However, professional ESG analysis is hindered by data fragmentation across unstructured sources, and existing large language models (LLMs) often struggle with the complex, multi-step workflows required for rigorous auditing. To address these limitations, we introduce ESGAgent, a hierarchical multi-agent system empowered by a specialized toolset, including retrieval augmentation, web search and domain-specific functions, to generate in-depth ESG analysis. Complementing this agentic system, we present a comprehensive three-level benchmark derived from 310 corporate sustainability reports, designed to evaluate capabilities ranging from atomic common-sense questions to the generation of integrated, in-depth analysis. Empirical evaluations demonstrate that ESGAgent outperforms state-of-the-art closed-source LLMs with an average accuracy of 84.15% on atomic question-answering tasks, and excels in professional report generation by integrating rich charts and verifiable references. These findings confirm the diagnostic value of our benchmark, establishing it as a vital testbed for assessing general and advanced agentic capabilities in high-stakes vertical domains.

ESG分析多智能体金融智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。