为大模型推理部署设计可持续性评估框架,支持跨地域数据中心的决策优化。
InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

- 基于查询轨迹和硬件模型构建可配置的能耗评估框架。
- 实测能耗估计误差小于10%,支持多数据中心扩展验证。
- 适合关注碳排放、能效与延迟权衡的基础设施规划者。
大模型推理的持续运行正将可持续性问题从训练阶段转向长期服务,基础设施选择直接影响能源消耗、碳排放、用水量和服务质量。然而,运营商在大规模建设前需评估多种部署方案,直接测量成本高、耗时长且不可行。本文提出InFactPlanner,一个基于追踪数据的决策支持框架,用于单点与地理分布式部署场景下的大模型推理可持续性“如果-那么”分析。该框架整合查询轨迹、硬件-模型性能画像、候选站点配置、PUE/WUE参数、可再生能源生成模型及随时间变化的电网碳强度,估算电力、能量、碳排放、用水量、延迟和服务器利用率。通过抽象底层服务影响为可配置的硬件-模型画像,实现对站点选择、容量分配、硬件选型、模型部署、可再生能源集成和路由策略的快速比较。我们通过重现参考大模型推理能耗估计,偏差低于10%验证了能源核算流程;评估了多数据中心与不同服务器数量下的可扩展性,并展示了针对硬件选型、可再生能源布局、地理分布和碳感知路由的场景化决策分析。结果表明,可持续性最优方案与延迟最优方案可能不同,且部署的碳效益高度依赖本地电网结构。
原文摘要 · Abstract (English)
The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to compare deployment alternatives before large-scale infrastructure is built, making direct measurement costly, slow, and sometimes infeasible. We present InFactPlanner, a trace-driven decision-support framework for what-if analysis of sustainable AI data center deployment for LLM inference across single and geo-distributed sites. InFactPlanner combines query traces, hardware-model profiles, candidate site configurations, PUE/WUE parameters, renewable generation models, and time-varying grid carbon intensity to estimate power, energy, carbon emissions, water use, latency, and server utilization. The framework abstracts low-level serving effects into configurable hardware-model profiles, enabling rapid comparison of site selection, capacity placement, hardware, model, renewable integration, and routing choices. We validate the energy accounting pipeline by reproducing reference LLM inference energy estimates with less than 10% deviation, evaluate scalability across multiple data centers and server counts, and demonstrate scenario-driven decision analyses for hardware selection, renewable placement, geographic deployment, and carbon-aware routing. Our results show that sustainability-optimal choices can differ from latency-optimal ones, and that the carbon value of deployment depends strongly on the local grid mix.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。