用大模型精准分析企业碳排放,解决知识滞后与报告解析难题
CarbonChat: Large Language Model-Based Corporate Carbon Emission Analysis and Climate Knowledge Q&A System
- 构建多维度指标模块,优化规则文档与长文本信息提取
- 设计增强型自提示检索生成架构,提升查询理解与转化效率
- 基于14维碳核算框架,支持报告总结与定制化问答,减少幻觉
随着全球气候变化影响加剧,企业碳排放成为关注焦点。针对大模型气候知识更新滞后、传统增强生成架构在复杂问题上缺乏专业性与准确性,以及可持续发展报告分析成本高、耗时长等问题,本文提出CarbonChat:基于大语言模型的企业碳排放分析与气候知识问答系统,旨在实现精准碳排放分析与政策理解。首先,提出多样化指标模块构建方法,用于处理基于规则与长文本文档的分割及结构化数据提取,优化关键信息解析。其次,设计增强型自提示检索-增强生成架构,融合意图识别、结构化推理链、混合检索与Text2SQL,提升语义理解与查询转换效率。再次,基于温室气体核算框架,建立14个维度的碳排放分析体系,支持报告摘要生成、相关性评估与定制化响应。最后,通过多层分块机制、时间戳记录与幻觉检测功能,确保分析结果的准确性和可验证性,降低幻觉率,提升回答精度。
原文摘要 · Abstract (English)
As the impact of global climate change intensifies, corporate carbon emissions have become a focal point of global attention. In response to issues such as the lag in climate change knowledge updates within large language models, the lack of specialization and accuracy in traditional augmented generation architectures for complex problems, and the high cost and time consumption of sustainability report analysis, this paper proposes CarbonChat: Large Language Model-based corporate carbon emission analysis and climate knowledge Q&A system, aimed at achieving precise carbon emission analysis and policy understanding.First, a diversified index module construction method is proposed to handle the segmentation of rule-based and long-text documents, as well as the extraction of structured data, thereby optimizing the parsing of key information.Second, an enhanced self-prompt retrieval-augmented generation architecture is designed, integrating intent recognition, structured reasoning chains, hybrid retrieval, and Text2SQL, improving the efficiency of semantic understanding and query conversion.Next, based on the greenhouse gas accounting framework, 14 dimensions are established for carbon emission analysis, enabling report summarization, relevance evaluation, and customized responses.Finally, through a multi-layer chunking mechanism, timestamps, and hallucination detection features, the accuracy and verifiability of the analysis results are ensured, reducing hallucination rates and enhancing the precision of the responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。