从企业报告中自动提取碳排放数据,且每条结果都有证据可查。
Scope3Trace: Evidence-Based Identification and Extraction of Scope 3 GHG Emissions from Sustainability Reports

- 用规则+大模型混合方式,精准定位并重建报告中的排放数据表。
- 在多源报告中准确提取组织与建筑级排放,类别级提取准确率高。
- 适合做碳足迹分析、环境合规审查的研究者和机构使用。
范围三温室气体(GHG)排放占企业碳足迹的大部分,但因披露稀疏、报告格式不一及证据追溯困难,难以规模化分析。现有方法多依赖大语言模型从ESG报告中提取信息,但常缺乏明确证据支撑,或需昂贵的人工标注验证可靠性。为此,我们提出Scope3Trace——一个基于证据的信息抽取框架,可从真实世界的企业可持续发展报告中提取可解释、可追溯的范围三排放数据。该框架包含文档收集与OCR解析、大模型辅助页面定位与表格重建、以及基于规则与大模型结合的组织与建筑层级排放抽取与证据验证流程。在此基础上,我们构建了一个双层级、基于证据的多模态数据集,涵盖来自异构报告的组织级范围三披露内容。Scope3Trace实现了对范围一至三总量及类别级排放的高精度提取,支持可靠且透明的跨报告整合。
原文摘要 · Abstract (English)
Scope 3 greenhouse gas (GHG) emissions account for the majority of corporate carbon footprints, yet remain difficult to analyze at scale due to sparse disclosures, heterogeneous report document formats, and limited evidence traceability. Existing approaches typically rely on large language models to extract emissions information from ESG reports, but often lack explicit evidence grounding or depend on costly manual annotation and verification to ensure extraction reliability. To address these challenges, we propose Scope3Trace, an evidence-grounded information extraction framework designed to extract interpretable and traceable Scope 3 emissions information from real-world ESG and sustainability reports. The framework integrates a document information extraction pipeline that performs PDF collection and OCR parsing, LLM-assisted page localization and table reconstruction, and hybrid rule-LLM extraction of organization- and building-level emissions disclosures with evidence-grounded verification. Building upon this framework, we further contribute a dual-level, evidence-grounded, multimodal dataset comprising organization-level Scope 3 disclosures extracted from heterogeneous sustainability reports. Scope3Trace enables reliable extraction and transparent integration of heterogeneous sustainability disclosures, achieving high accuracy in extracting Scope 1-3 totals and category-level disclosures from sustainability reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。