arXiv:2510.26136cs.AI2025-10被引 3

量化大模型推理成本,揭示其经济规律与最优部署区间。

Beyond Benchmarks: The Economics of AI Inference

  • 将大模型推理视为计算驱动的生产活动,构建经济分析框架。
  • 基于WiNEval-3.0数据,发现边际成本递减、规模收益递减规律。
  • 提出推理效率最优区,指导模型部署与资源定价决策。

大型语言模型(LLMs)的推理成本已成为决定其商业可行性和广泛应用的关键因素。本文提出一个量化的「推理经济学」框架,将LLM推理过程视为由计算驱动的智能生产活动。我们分析了不同性能配置下的边际成本、规模经济性及输出质量。基于WiNEval-3.0的实证数据,首次构建了「LLM推理生产前沿」,揭示三大规律:边际成本递减、规模收益递减,以及最优成本效益区间。该研究不仅为模型部署提供经济依据,也为未来基于市场的推理资源定价与优化奠定实证基础。

原文摘要 · Abstract (English)

The inference cost of Large Language Models (LLMs) has become a critical factor in determining their commercial viability and widespread adoption. This paper introduces a quantitative ``economics of inference'' framework, treating the LLM inference process as a compute-driven intelligent production activity. We analyze its marginal cost, economies of scale, and quality of output under various performance configurations. Based on empirical data from WiNEval-3.0, we construct the first ``LLM Inference Production Frontier,'' revealing three principles: diminishing marginal cost, diminishing returns to scale, and an optimal cost-effectiveness zone. This paper not only provides an economic basis for model deployment decisions but also lays an empirical foundation for the future market-based pricing and optimization of AI inference resources.

推理成本大模型经济学部署优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。