让大模型在边缘设备上更省电低碳,还能保持高效运行。
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
- 动态选择工具并根据碳强度调整能耗阈值
- 比传统方法减排52%,降耗30%,提速30%
- 适合注重环保与能效的边缘AI部署
大型语言模型(LLMs)虽能在边缘AI系统中实现实时函数调用,但带来显著计算开销,导致高功耗和碳排放。现有方法侧重性能优化,忽视可持续性,难以适应能源受限环境。我们提出CarbonCall,一种面向可持续性的函数调用框架,融合动态工具选择、碳感知执行与量化LLM适配。CarbonCall基于实时碳强度预测调整功耗阈值,并在不同模型版本间切换,以在功耗约束下维持高每秒词元吞吐量。在NVIDIA Jetson AGX Orin上的实验表明,CarbonCall可使碳排放降低最多52%,功耗减少30%,执行时间缩短30%,同时保持高效率。
原文摘要 · Abstract (English)
Large Language Models (LLMs) enable real-time function calling in edge AI systems but introduce significant computational overhead, leading to high power consumption and carbon emissions. Existing methods optimize for performance while neglecting sustainability, making them inefficient for energy-constrained environments. We introduce CarbonCall, a sustainability-aware function-calling framework that integrates dynamic tool selection, carbon-aware execution, and quantized LLM adaptation. CarbonCall adjusts power thresholds based on real-time carbon intensity forecasts and switches between model variants to sustain high tokens-per-second throughput under power constraints. Experiments on an NVIDIA Jetson AGX Orin show that CarbonCall reduces carbon emissions by up to 52%, power consumption by 30%, and execution time by 30%, while maintaining high efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。