arXiv:2604.27855cs.DCcs.AI2026-04

让AI推理像电力一样可迁移,按延迟限制优化能源使用。

AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography Framework

  • 构建三层架构模型,统筹延迟、碳排与算力分配。
  • 发现延迟放宽可拓展计算地理范围,提升能源效率。
  • 适合关注绿色计算与云调度的工程师和政策制定者。

AI推理正成为持续且分布式的电力需求源。与传统用电负载不同,只要延迟、状态本地性、容量及监管约束可接受,推理任务可远离用户服务位置执行。本文研究如何将此类数字计算迁移视为延迟约束下的电力需求迁移。提出一种面向地理分布式AI推理的能量-地理框架,建模客户端、服务节点与计算节点的三层结构,将推理部署建模为在电价、边际碳强度、电源使用效率、算力容量、网络延迟与迁移摩擦约束下的优化问题。核心是能量-延迟前沿:放松推理延迟预算所释放的边际成本降低与碳减排收益。论文贡献包括:区分物理电能传输与计算任务的数字迁移;引入可行性掩码与迁移摩擦的部署模型;提出可迁移推理需求、能量延迟回报、碳延迟回报及迁移盈亏平衡等运营指标;通过典型全球算力区域的透明模拟,揭示异构延迟容忍度使工作负载分层为本地、区域与能源导向执行。结果表明,延迟放宽可扩大可行地理范围,但迁移摩擦、出站成本、状态本地性、法律限制与容量瓶颈会显著削弱实际效益。

原文摘要 · Abstract (English)

AI inference is becoming a persistent and geographically distributed source of electricity demand. Unlike many traditional electrical loads, inference workloads can sometimes be executed away from the user-facing service location, provided that latency, state locality, capacity, and regulatory constraints remain acceptable. This paper studies when such digital relocation of computation can be interpreted as latency-constrained relocation of electricity demand. We develop an energy-geography framework for geo-distributed AI inference. The framework models a three-layer architecture of clients, service nodes, and compute nodes, and formulates inference placement as a constrained optimization problem over electricity prices, marginal carbon intensity, power usage effectiveness, compute capacity, network latency, and migration frictions. The key object is the energy-latency frontier: the marginal cost and carbon benefit unlocked by relaxing inference latency budgets. The paper makes four contributions. First, it distinguishes physical electricity transmission from digital relocation of electricity-consuming computation. Second, it formulates a geo-distributed inference placement model with feasibility masks and migration frictions. Third, it introduces operational metrics, including relocatable inference demand, energy return on latency, carbon return on latency, and a relocation break-even condition. Fourth, it provides a transparent stylized simulation over representative global compute regions to show how heterogeneous latency tolerance separates workloads into local, regional, and energy-oriented execution layers. The results show that latency relaxation expands feasible geography, while migration frictions, egress costs, state locality, legal constraints, and capacity limits can sharply reduce realized benefits.

AI推理能源优化地理调度碳排放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。