arXiv:2605.22883cs.AIcs.LG2026-05被引 1

为智能体系统设计新能效指标,按成功目标计耗电。

Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems

论文配图:Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems
图 1 · 摘自论文原文
  • 用‘每成功目标耗能’替代‘每推理耗能’作为计量单位。
  • 智能体系统平均能耗是线性执行的4.33倍,达888.1焦耳/成功目标。
  • 揭示编排结构才是能耗主因,适合评估智能体系统能效。

当前AI能效基准以单次模型调用或训练运行为单位,对单轮任务仍合理。但对智能体系统——一个用户目标可能引发多步编排、工具调用、重试与失败恢复——调用次数仅为实现细节而非任务属性,以推理级归一化会扭曲目标完成的真实能耗。本文提出A-LEMS(智能体大模型能效测量系统),构建跨层测量框架,将能效计量单位从‘每推理耗能’重构为‘每成功目标耗能’(EpG),聚合所有执行尝试(含失败与重试)的总能耗,并按成功目标数归一化。A-LEMS通过时间边界模型、五层观测流水线(映射RAPL信号至工作流级能耗)及可复现协议,确保测量绑定硬件与运行环境配置。基于EpG,定义编排开销指数(OOI),隔离编排能耗相对于线性执行的成本。在五个推理与三个工具增强任务族中,智能体工作流平均能耗为888.1焦耳/成功目标,是线性基线205.3焦耳的4.33倍。该开销源于编排结构而非推理计算本身。对于工具增强任务,当OOI低于1.0时,智能体执行反而更省电,验证了该指标捕捉编排结构而非固定上偏。结论表明:‘每推理耗能’不足以衡量智能体系统能效。EpG与OOI为准确基准提供基础,其中编排结构是能耗的核心决定因素。

原文摘要 · Abstract (English)

Current AI energy benchmarks measure consumption at the granularity of a single model invocation or training run. For classical single-turn workloads this unit remains coherent. For agentic systems - where a single user goal may trigger multi-step orchestration, tool calls, retries, and failure-recovery cycles - the invocation count is an implementation artifact rather than a task property, and inference-level normalization misrepresents the energy cost of goal completion. We present A-LEMS (Agentic LLM Energy Measurement System), a cross-layer measurement framework that redefines the unit of AI energy accounting from energy per inference to Energy per Successful Goal (EpG). EpG aggregates total workflow energy across all execution attempts, including failures and retries, normalized by successfully completed goals. A-LEMS formalizes energy attribution through a temporal boundary model, a five-layer observation pipeline mapping RAPL signals to workflow-level energy, and a reproducibility protocol binding every measurement to hardware and runtime configuration. Building on EpG, we define the Orchestration Overhead Index (OOI), isolating the energy cost of orchestration relative to linear execution under identical task criteria. Across five reasoning and three tool-augmented task families, agentic workflows consume 4.33x higher mean energy per successful goal than linear baselines (888.1 J vs 205.3 J). This overhead is driven by orchestration structure, not inference compute. For tool-augmented tasks, OOI inverts below 1.0x: agentic execution is cheaper than linear, confirming the metric captures orchestration structure rather than a fixed upward bias. These findings establish that energy-per-inference is insufficient for agentic AI. EpG and OOI provide the measurement foundation for accurate benchmarking, where orchestration structure is the primary determinant of energy cost.

能效评估智能体系统编排开销

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。