让本地运行大模型更省电、更透明,还能自动推荐优化方案。
EnviroLLM: Resource Tracking and Optimization for Local AI
- 实时监控本地大模型的能耗与性能,支持多平台对比
- 结合能效与速度评估,量化模型在不同场景下的表现
- 适合关注隐私、能效或想优化本地部署的开发者
大型语言模型(LLMs)正越来越多地在本地部署以保障隐私和可访问性,但用户缺乏工具来衡量其资源消耗、环境影响及效率指标。本文提出 EnviroLLM,一个开源工具包,用于追踪、基准测试和优化在个人设备上运行 LLM 时的性能与能耗。系统提供实时进程监控,支持 Ollama、LM Studio、vLLM 及 OpenAI 兼容接口等多平台基准测试,具备持久化存储与可视化功能,支持长期分析,并给出个性化模型与优化建议。系统还引入大模型作为评判者,结合能耗与速度指标,帮助用户在使用自定义提示词测试模型时评估质量-效率权衡。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed locally for privacy and accessibility, yet users lack tools to measure their resource usage, environmental impact, and efficiency metrics. This paper presents EnviroLLM, an open-source toolkit for tracking, benchmarking, and optimizing performance and energy consumption when running LLMs on personal devices. The system provides real-time process monitoring, benchmarking across multiple platforms (Ollama, LM Studio, vLLM, and OpenAI-compatible APIs), persistent storage with visualizations for longitudinal analysis, and personalized model and optimization recommendations. The system includes LLM-as-judge evaluations alongside energy and speed metrics, enabling users to assess quality-efficiency tradeoffs when testing models with custom prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。