arXiv:2608.27716cs.AI2026-08

首个针对产品碳足迹估算的诊断基准,揭示大模型在分解推理中的系统性缺陷。

PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation

论文配图:PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation
图 1 · 摘自论文原文
  • 将碳足迹估算拆解为6个可独立评估的任务,涵盖分解、检索、匹配与数值提取
  • 8个主流大模型在分步推理中仅37%-58%准确,质量显著低于整体结果
  • 暴露模型在质量守恒等约束下的漏洞,适合关注可信低碳决策的研究者

AI系统正被应用于高风险、领域特定的工作流,要求不仅最终输出正确,中间步骤也需准确。以产品碳足迹(PCF)估算为例,即物理产品所关联的温室气体排放量。尽管越来越多的AI代理被用于生成PCF,但现有评估要么仅评分总排放量(掩盖错误来源并抵消误差),要么孤立评估子任务(忽略组合交互)。我们提出PCFBench,首个将PCF建模分解为可独立评估任务的基准,涵盖分解、信息检索、本体匹配和数值提取。该数据集包含614个专家标注项,覆盖6个任务,能探测在信息不全、上下文冲突及数值约束下的推理能力。在来自四个供应商的8个前沿大模型中,无一模型全面领先。最强模型虽能在77%的产品上使总排放估计值接近声明值(误差在2倍以内),但在分步生成时准确率降至37%-58%,且仅有45%-75%满足质量守恒。这些失败削弱了从业者进行产品对比和推动脱碳所需的透明度。我们已公开数据集与评估工具,以支持针对性改进。

原文摘要 · Abstract (English)

AI systems are being deployed on high-stakes, domain-specific workflows that demand correctness not just in the final output, but at every intermediate step. One such workflow is estimating a product carbon footprint (PCF), the greenhouse-gas emissions attributable to a physical product. AI agents are increasingly being used to generate PCFs, but existing evaluations score either total emissions (hiding error sources and cancelling mistakes) or sub-tasks in isolation (missing compositional interactions). We introduce PCFBench, the first benchmark to carve PCF modeling into independently-evaluable tasks that require decomposition, retrieval, ontology matching, and numerical extraction. It comprises 614 expert-labelled items across six tasks. Together they probe reasoning under under-specification, conflicting context, and numerical constraints. Across eight frontier LLMs from four providers, no single model dominates. Although the strongest models estimate total product emissions within 2 times of declared totals on 77% of products, this rate drops to 37-58% when the PCF is generated step by step, with only 45-75% obeying mass conservation. These failures undermine the transparency practitioners need to compare products and drive decarbonization. We release the dataset and evaluation harness to support targeted progress.

碳足迹大模型评测可解释性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。