用AI统一企业碳排放估算,解决中小企业难算碳的问题。
Group Reasoning Emission Estimation Networks
- 将行业分类转为信息检索,用对比学习优化文本匹配
- 在1114个行业上达83.68%准确率,20家企业平均误差45.88%
- 基于NAICS层级结构分步推理,适合需精准碳核算的用户
准确的温室气体排放报告对政府、企业和投资者至关重要。然而,由于实施成本高、排放因子数据库分散及缺乏稳健的行业分类方法,中小企业采纳率仍低。为此,我们提出群组推理排放估算网络(GREEN),一种基于AI的企业级碳核算框架,标准化排放估算流程,构建大规模基准数据集,并利用大语言模型实现新型推理方法。具体而言,我们收集了20,850家企业的文本描述,并与经验证的北美行业分类系统(NAICS)标签对齐,结合碳强度经济模型。通过将行业分类重构为信息检索任务,采用对比学习损失微调Sentence-BERT模型。针对千级层级分类中单阶段模型的局限性,提出基于自然NAICS本体的群组推理方法,将任务分解为多步子分类,理论上降低分类不确定性和计算复杂度。在1,114个NAICS类别上的实验达到最优性能(Top-1准确率83.68%,Top-10准确率91.47%),20家企业的案例研究显示平均绝对百分比误差(MAPE)为45.88%。项目已开源:https://huggingface.co/datasets/Yvnminc/ExioNAICS。
原文摘要 · Abstract (English)
Accurate greenhouse gas (GHG) emission reporting is critical for governments, businesses, and investors. However, adoption remains limited particularly among small and medium enterprises due to high implementation costs, fragmented emission factor databases, and a lack of robust sector classification methods. To address these challenges, we introduce Group Reasoning Emission Estimation Networks (GREEN), an AI-driven carbon accounting framework that standardizes enterprise-level emission estimation, constructs a large-scale benchmark dataset, and leverages a novel reasoning approach with large language models (LLMs). Specifically, we compile textual descriptions for 20,850 companies with validated North American Industry Classification System (NAICS) labels and align these with an economic model of carbon intensity factors. By reframing sector classification as an information retrieval task, we fine-tune Sentence-BERT models using a contrastive learning loss. To overcome the limitations of single-stage models in handling thousands of hierarchical categories, we propose a Group Reasoning method that ensembles LLM classifiers based on the natural NAICS ontology, decomposing the task into multiple sub-classification steps. We theoretically prove that this approach reduces classification uncertainty and computational complexity. Experiments on 1,114 NAICS categories yield state-of-the-art performance (83.68% Top-1, 91.47% Top-10 accuracy), and case studies on 20 companies report a mean absolute percentage error (MAPE) of 45.88%. The project is available at: https://huggingface.co/datasets/Yvnminc/ExioNAICS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。