Bench360统一评估本地大模型推理性能,帮用户选最优配置。
Bench360: Benchmarking Local LLM Inference from 360 Degrees
- 整合多引擎、量化格式与任务,统一评测本地大模型推理
- 在3张GPU上测试4类任务,揭示配置对效率与质量的显著影响
- 适合关注实际部署的开发者与研究人员
本地运行大语言模型日益普遍,但用户面临模型、量化级别、推理引擎和部署场景的复杂选择。现有基准测试分散且聚焦单一目标,难以指导实际部署。我们提出Bench360,一个集成多任务、使用模式与系统指标的本地LLM推理评估框架。该框架支持自定义任务,集成多个推理引擎与量化格式,并报告任务质量及系统行为(延迟、吞吐、能耗、启动时间)。我们在三张GPU上对四类NLP任务进行了验证,展示了不同设计选择如何影响效率与输出质量。结果表明,权衡关系显著,最优配置依赖具体工作负载与约束条件。不存在通用最优方案,凸显了综合性、面向部署的基准测试的必要性。
原文摘要 · Abstract (English)
Running LLMs locally has become increasingly common, but users face a complex design space across models, quantization levels, inference engines, and serving scenarios. Existing inference benchmarks are fragmented and focus on isolated goals, offering little guidance for practical deployments. We present Bench360, a framework for evaluating local LLM inference across tasks, usage patterns, and system metrics in one place. Bench360 supports custom tasks, integrates multiple inference engines and quantization formats, and reports both task quality and system behavior (latency, throughput, energy, startup time). We demonstrate it on four NLP tasks across three GPUs and four engines, showing how design choices shape efficiency and output quality. Results confirm that tradeoffs are substantial and configuration choices depend on specific workloads and constraints. There is no universal best option, underscoring the need for comprehensive, deployment-oriented benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。