为大模型驱动的软件构建高效性能工程体系
Software Performance Engineering for Foundation Model-Powered Software
- 提出大模型软件性能工程四重挑战:架构设计、通信协议、调优优化与部署
- 强调性能需前置规划,避免上线后高成本重构
- 适合大模型应用开发与系统架构师参考
大模型(如LLM)正重塑软件开发,但将原型转化为生产级产品需跨领域复杂工程。其中性能工程常被忽视,直接影响吞吐量与延迟,导致用户不满与经济损失。当前多数性能考量滞后,造成部署后昂贵优化。大模型对算力需求高,亟需高效硬件利用。持续性能工程可防止性能退化。本文基于文献与自研系统经验,提出四类核心挑战:认知架构设计(即AI组件间交互、推理与传统软件接口的结构设计)、通信协议、调优优化与部署。分析现存问题与实践,探索软件工程界创新路径。
原文摘要 · Abstract (English)
The rise of Foundation Models (FMs) like Large Language Models (LLMs) is revolutionizing software development. Despite the impressive prototypes, transforming FMware into production-ready products demands complex engineering across various domains. A critical but overlooked aspect is performance engineering, which aims at ensuring FMware meets performance goals such as throughput and latency to avoid user dissatisfaction and financial loss. Often, performance considerations are an afterthought, leading to costly optimization efforts post-deployment. FMware's high computational resource demands highlight the need for efficient hardware use. Continuous performance engineering is essential to prevent degradation. This paper highlights the significance of Software Performance Engineering (SPE) in FMware, identifying four key challenges: cognitive architecture design (i.e., the structural design that defines how AI components interact, reason, and interface with classical software components), communication protocols, tuning and optimization, and deployment. These challenges are based on literature surveys and experiences from developing an in-house FMware system. We discuss problems, current practices, and innovative paths for the software engineering community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。