arXiv:2510.27628cs.AI2025-10

Agentic AI的成功关键在于用户验证,而非依赖大模型。

Validity Is What You Need

  • 将智能体定义为自主运行的软件交付方式,强调实际应用价值。
  • 真实验证可替代复杂大模型,用简单模型也能实现核心功能。
  • 适合关注系统落地与可信度的企业级开发者和产品经理。

尽管人工智能代理在计算机科学中长期被讨论,但当前的智能体人工智能系统是全新的。本文提出一种新的现实主义定义:智能体AI是一种软件交付机制,类似于软件即服务(SaaS),可在复杂的组织环境中自主运行应用程序。大型语言模型(LLMs)作为基础模型的进展推动了智能体AI的发展,但本文指出,智能体系统本质上是应用,而非基础,其成功取决于最终用户和主要利益相关者的验证。用于验证应用的工具与评估基础模型的工具截然不同。讽刺的是,有了良好的验证机制后,许多情况下可用更简单、更快、更可解释的模型替代基础模型来处理核心逻辑。对于智能体AI而言,真正需要的是有效性。大语言模型只是可能达成有效性的选项之一。

原文摘要 · Abstract (English)

While AI agents have long been discussed and studied in computer science, today's Agentic AI systems are something new. We consider other definitions of Agentic AI and propose a new realist definition. Agentic AI is a software delivery mechanism, comparable to software as a service (SaaS), which puts an application to work autonomously in a complex enterprise setting. Recent advances in large language models (LLMs) as foundation models have driven excitement in Agentic AI. We note, however, that Agentic AI systems are primarily applications, not foundations, and so their success depends on validation by end users and principal stakeholders. The tools and techniques needed by the principal users to validate their applications are quite different from the tools and techniques used to evaluate foundation models. Ironically, with good validation measures in place, in many cases the foundation models can be replaced with much simpler, faster, and more interpretable models that handle core logic. When it comes to Agentic AI, validity is what you need. LLMs are one option that might achieve it.

智能体AI有效性验证大模型替代企业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。