用用户激励降低AI推理碳排放,兼顾质量与延迟
Greening AI Inference with Accuracy and Latency-aware User Incentives

- 根据用户对质量、延迟和环保的偏好设计分层激励
- 通过折扣服务让部分请求降质延时,降低碳排放
- 适合关注绿色AI与用户体验平衡的研究者
AI服务的广泛应用引发环境可持续性担忧,现有研究指出推理阶段的碳排放是主要来源。本文提出一种基于用户对推理质量与延迟的偏好及环保意识的推理激励框架,同时考虑碳排放与服务质量(QoE)参数之间的权衡。该方法可适应不同模型规模与资源分配下的多种权衡策略。通过实用的两级订阅服务,用户以降低服务质量与增加延迟为代价换取碳排放减免的折扣。在高碳强度时段,服务商可灵活地以较低质量、较高延迟处理部分推理请求,实现低碳运行。
原文摘要 · Abstract (English)
The widespread use of AI services has raised concerns for its environmental sustainability, towards which recent studies have identified carbon emissions of AI inference as the major contributor. This paper introduces a framework for designing AI inference incentives based on the users' valuation for inference quality and latency, together with their environmental consciousness, while accounting for the tradeoff between carbon emissions and the two QoE parameters. Our approach can accommodate different tradeoffs, that depend on the size and complexity of the AI models and the allocation of resources to serve inference requests. The incentives can be offered through a practical two-tier service subscription that offers users a discount in exchange for reduced carbon emissions. The discounted service option gives the AI provider the flexibility to serve some percentage of inference requests at a lower quality and higher latency during periods of high carbon intensity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。