首次系统研究生产环境中的大模型代理,揭示其实际构建与评估方式。
Measuring Agents in Production
- 基于20个深度访谈和86名实践者调研,分析真实部署场景。
- 超六成代理在10步内需人工介入,七成依赖现成模型提示而非调参。
- 可靠性是核心挑战,当前通过系统设计应对,适合关注落地的开发者参考。
基于大语言模型的智能体已在多个行业投入生产,但我们对使部署成功的技术方法仍缺乏理解。本文首次开展生产环境中智能体的系统性研究——测量生产中的智能体(MAP),基于一线开发者的实际数据。通过20个深度访谈及覆盖26个领域的86名已部署系统实践者的调查,我们探究组织构建智能体的原因、构建方式、评估策略及主要开发挑战。研究发现,生产级智能体普遍采用简单可控的方法:68%的智能体在人类介入前最多执行10步,70%依赖提示现成模型而非权重调优,74%主要依靠人工评估。可靠性(长时间保持正确行为的一致性)仍是首要挑战,从业者目前通过系统级设计来应对。MAP 揭示了当前生产智能体的真实状态,为研究社区提供了对部署现实的洞见,并指出了尚未充分探索的研究方向。
原文摘要 · Abstract (English)
LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first systematic study of Measuring Agents in Production, MAP, using first-hand data from agent developers. We conducted 20 case studies via in-depth interviews and surveyed 86 deployed systems practitioners across 26 domains. We investigate why organizations build agents, how they build them, how they evaluate them, and their top development challenges. Our study finds that production agents are built using simple, controllable approaches: 68% execute at most 10 steps before human intervention, 70% rely on prompting off-the-shelf models instead of weight tuning, and 74% depend primarily on human evaluation. Reliability (consistent correct behavior over time) remains the top development challenge, which practitioners currently address through systems-level design. MAP documents the current state of production agents, providing the research community with visibility into deployment realities and underexplored research avenues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。