arXiv:2606.20295cs.SEcs.CL2026-06

提出四层架构优化大模型推理,降低令牌成本,提升服务稳定性。

Token-Operations-Oriented Inference Optimization Techniques for Large Models

论文配图:Token-Operations-Oriented Inference Optimization Techniques for Large Models
图 1 · 摘自论文原文
  • 构建多层级融合架构,聚焦令牌层面优化
  • 实现令牌生产成本下降与服务效率提升
  • 适合追求高效稳定大模型部署的工程团队

大模型推理优化是支撑大模型服务可扩展、低成本、高稳定运行的关键基础。本文首次提出以令牌为导向的四层技术架构:多模型融合、模型优化、计算-模型融合、计算-网络-模型融合。系统梳理了各层级的关键技术与行业现状,分析了相关技术在真实业务场景中的应用价值。该研究为降低令牌生成成本、提升令牌服务效率、保障令牌供应稳定性提供了可行的技术路径,推动大模型服务从仅可调用向可运维转型。

原文摘要 · Abstract (English)

Large model inference optimization serves as a key foundation for supporting the scalable, low-cost, and highly stable operation of large model services. Centered on token-oriented inference optimization technology, this paper proposes for the first time a four-layer technical architecture consisting of Multi-model Fusion, Model Optimization, Compute-Model Fusion, and Compute-Network-Model Fusion. It systematically reviews the key technologies and current industry status across these four levels and analyzes the application value of related technologies in real-world business scenarios. This paper provides a practical technical path for reducing token production costs, improving token service efficiency, ensuring the stability of token supply, and driving the transition of large model services from being merely callable to being operable.

大模型优化推理加速令牌管理架构设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。