小模型更适合重复性任务,省钱又高效。
Small Language Models are the Future of Agentic AI
- 用小模型替代大模型完成特定任务
- 部署成本降低,响应速度更快
- 适合追求性价比的智能体系统
大型语言模型(LLMs)在多种任务中表现出接近人类的能力,擅长通用对话。然而,智能体系统(agentic AI)的兴起带来了大量重复性、低变异性的专项任务。本文主张:小语言模型(SLMs)已具备足够能力,天生更适配此类场景,且在部署上更具经济性,因此将是未来智能体AI的核心。该观点基于当前SLMs的能力水平、智能体系统的常见架构以及模型部署的经济性。我们进一步指出,若需通用对话能力,应采用异构智能体系统(即调用多个不同模型)。文中讨论了推广SLMs的潜在障碍,并提出一个通用的从LLM到SLM的智能体转换算法。我们的核心主张强调:即使部分转向SLMs,也能对智能体产业带来显著的运营与经济影响。旨在推动对AI资源高效利用的讨论,助力降低当前人工智能的成本。欢迎各界对本立场进行贡献或批评,所有回应将发布于 https://research.nvidia.com/labs/lpr/slm-agents。
原文摘要 · Abstract (English)
Large language models (LLMs) are often praised for exhibiting near-human performance on a wide range of tasks and valued for their ability to hold a general conversation. The rise of agentic AI systems is, however, ushering in a mass of applications in which language models perform a small number of specialized tasks repetitively and with little variation. Here we lay out the position that small language models (SLMs) are sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems, and are therefore the future of agentic AI. Our argumentation is grounded in the current level of capabilities exhibited by SLMs, the common architectures of agentic systems, and the economy of LM deployment. We further argue that in situations where general-purpose conversational abilities are essential, heterogeneous agentic systems (i.e., agents invoking multiple different models) are the natural choice. We discuss the potential barriers for the adoption of SLMs in agentic systems and outline a general LLM-to-SLM agent conversion algorithm. Our position, formulated as a value statement, highlights the significance of the operational and economic impact even a partial shift from LLMs to SLMs is to have on the AI agent industry. We aim to stimulate the discussion on the effective use of AI resources and hope to advance the efforts to lower the costs of AI of the present day. Calling for both contributions to and critique of our position, we commit to publishing all such correspondence at https://research.nvidia.com/labs/lpr/slm-agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。