为大模型系统构建可复用组件,需补上规格说明书这一关键短板。
Specifications: The missing link to making the development of LLM systems an engineering discipline
- 提出用精确规格定义大模型组件行为,提升系统可组合性。
- 现有方法如结构化输出、过程监督等已初见成效。
- 适合关注大模型工程化落地的研究者与开发者。
尽管生成式AI在短短几年内取得显著进展,但其未来发展受限于构建模块化、可靠系统的能力。历史上,汽车、飞机、计算机和软件的成功都依赖于可组装、可调试、可替换的组件(如发动机、轮子、CPU、库)。实现这一目标的关键工具是规格说明:对组件预期行为、输入和输出的精确描述。然而,大模型的通用性和自然语言的模糊性,使得为大模型组件(如智能体)制定规格成为一项极具挑战且紧迫的问题。本文讨论了当前领域在结构化输出、过程监督、测试时计算等方面的进展,并指出了未来研究方向,旨在通过改进规格说明,推动大模型系统向工程化发展。
原文摘要 · Abstract (English)
Despite the significant strides made by generative AI in just a few short years, its future progress is constrained by the challenge of building modular and robust systems. This capability has been a cornerstone of past technological revolutions, which relied on combining components to create increasingly sophisticated and reliable systems. Cars, airplanes, computers, and software consist of components-such as engines, wheels, CPUs, and libraries-that can be assembled, debugged, and replaced. A key tool for building such reliable and modular systems is specification: the precise description of the expected behavior, inputs, and outputs of each component. However, the generality of LLMs and the inherent ambiguity of natural language make defining specifications for LLM-based components (e.g., agents) both a challenging and urgent problem. In this paper, we discuss the progress the field has made so far-through advances like structured outputs, process supervision, and test-time compute-and outline several future directions for research to enable the development of modular and reliable LLM-based systems through improved specifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。