提出一套统一的AI模型推理评估方法,兼顾性能与碳排放。
Metrics and evaluations for computational and sustainable AI efficiency
- 整合延迟、吞吐、能耗与碳排放,实现多维度评估
- 在多种硬件和精度下验证,发现能效与碳排存在显著差异
- 开源工具支持复现,助力可持续AI部署决策
人工智能的快速发展带来了巨大的算力需求,但现有模型性能、效率与环境影响的评估方法仍零散不一。当前方法难以全面比较不同硬件、软件栈和数值精度下的系统表现。为此,本文提出一种统一且可复现的AI模型推理评估方法,集成计算与环境指标,在真实服务条件下系统测量延迟、吞吐分布、能耗及位置相关的碳排放,同时保持相同的准确率约束以确保可比性。该框架应用于从数据中心加速器GH200到消费级显卡RTX 4090等多种硬件平台上的多精度模型,运行于PyTorch、TensorRT和ONNX Runtime等主流软件栈。通过系统分类各因素,建立严谨的基准测试体系,生成决策可用的帕累托前沿,明确揭示准确率、延迟、能耗与碳排放之间的权衡关系。配套开源代码支持独立验证与推广,帮助研究者与实践者基于证据做出可持续的AI部署选择。
原文摘要 · Abstract (English)
The rapid advancement of Artificial Intelligence (AI) has created unprecedented demands for computational power, yet methods for evaluating the performance, efficiency, and environmental impact of deployed models remain fragmented. Current approaches often fail to provide a holistic view, making it difficult to compare and optimise systems across heterogeneous hardware, software stacks, and numeric precisions. To address this gap, we propose a unified and reproducible methodology for AI model inference that integrates computational and environmental metrics under realistic serving conditions. Our framework provides a pragmatic, carbon-aware evaluation by systematically measuring latency and throughput distributions, energy consumption, and location-adjusted carbon emissions, all while maintaining matched accuracy constraints for valid comparisons. We apply this methodology to multi-precision models across diverse hardware platforms, from data-centre accelerators like the GH200 to consumer-level GPUs such as the RTX 4090, running on mainstream software stacks including PyTorch, TensorRT, and ONNX Runtime. By systematically categorising these factors, our work establishes a rigorous benchmarking framework that produces decision-ready Pareto frontiers, clarifying the trade-offs between accuracy, latency, energy, and carbon. The accompanying open-source code enables independent verification and facilitates adoption, empowering researchers and practitioners to make evidence-based decisions for sustainable AI deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。