arXiv:2512.23029cs.DCcs.AI2025-12

用消费级显卡部署私有大模型,让中小企业低成本用上高性能AI。

Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware

  • 在消费级硬件上运行量化后的Qwen3-30B MoE模型,实现本地化部署。
  • 延迟低至350ms,每秒生成超过40个词元,支持多用户并发。
  • 适合关注数据隐私与成本控制的中小企业技术决策者。

大型语言模型的普及依赖于云服务,带来数据隐私、运营主权和成本上升等挑战。本文研究了中小企业在可负担成本下部署高性能私有LLM推理服务的可行性。我们对基于Qwen3的300亿参数混合专家(MoE)模型进行了全面基准测试,该模型在配备新一代NVIDIA GPU的消费级服务器上运行并经过量化处理。相比昂贵且复杂的云端方案,本方法为中小企业提供了低成本、私密的解决方案。评估涵盖模型能力与服务器性能两方面:模型表现通过学术与行业标准对比,衡量其推理与知识水平;服务器效率则以延迟、每秒词元数及首词元响应时间衡量,并分析并发用户增加时的可扩展性。结果表明,合理配置的本地部署系统在新兴消费级硬件上可达到与云端服务相当的性能,为中小企业提供无需高昂成本或隐私风险即可使用强大LLM的可行路径。

原文摘要 · Abstract (English)

The proliferation of Large Language Models (LLMs) has been accompanied by a reliance on cloud-based, proprietary systems, raising significant concerns regarding data privacy, operational sovereignty, and escalating costs. This paper investigates the feasibility of deploying a high-performance, private LLM inference server at a cost accessible to Small and Medium Businesses (SMBs). We present a comprehensive benchmarking analysis of a locally hosted, quantized 30-billion parameter Mixture-of-Experts (MoE) model based on Qwen3, running on a consumer-grade server equipped with a next-generation NVIDIA GPU. Unlike cloud-based offerings, which are expensive and complex to integrate, our approach provides an affordable and private solution for SMBs. We evaluate two dimensions: the model's intrinsic capabilities and the server's performance under load. Model performance is benchmarked against academic and industry standards to quantify reasoning and knowledge relative to cloud services. Concurrently, we measure server efficiency through latency, tokens per second, and time to first token, analyzing scalability under increasing concurrent users. Our findings demonstrate that a carefully configured on-premises setup with emerging consumer hardware and a quantized open-source model can achieve performance comparable to cloud-based services, offering SMBs a viable pathway to deploy powerful LLMs without prohibitive costs or privacy compromises.

私有部署中小企业量化模型本地推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。