构建千万级用户客服AI Agent的评估驱动框架,实现快速迭代与高满意度。
Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework

- 构建端到端评估框架,整合上下文工程与人工反馈优化提示。
- 卡交付场景下,自服务率提升29个百分点,客户满意度提高37个百分点。
- 离线评估指标与线上结果高度相关,可精准预测实际效果。
大模型能力的快速发展使AI代理在众多任务中日益可行,尤其在构建面向客户的生产级客服代理方面前景广阔。然而,这需要评估方法、上下文工程、训练和在线度量等多方面的协同突破,而这些关键环节常被孤立开发,导致部署后才暴露问题。本文提出一个统一框架,将离线开发与线上影响相衔接,应用于拥有超1亿用户的Nubank公司。该框架包含:(1)针对客服场景的结构化上下文工程;(2)系统化的人工参与提示迭代;(3)严格的LLM评判机制,通过测量评分者间一致性并优化GEPA以保证一致性;(4)从构思到上线的全流程验证。核心发现是:评估管道质量直接决定迭代速度。五次不同领域部署(卡配送、债务管理、信用额度支持、卡片管理、产品说明)均实现客户满意度持续提升,并显著加速迭代。在卡配送场景中,大规模A/B测试显示,AI交易型净推荐值提升37个百分点,自服务率提升29个百分点,且离线模拟指标与线上结果高度相关,证明评估驱动开发能可靠预测生产影响。多数场景下,AI满意度已接近专家人类客服水平。
原文摘要 · Abstract (English)
The rapid rise in LLM capabilities has made AI agents increasingly viable across a broad range of tasks. Among the most promising applications is building production-ready customer-facing agents, a challenge that demands coordinated excellence in evaluation methodology, context engineering, training, and online measurement. Yet these critical pillars are typically developed in isolation, creating blind spots that only surface after deployment. In this paper, we present a unified framework that bridges offline development with online impact for customer support AI agents at Nubank, a company with 100M+ users. Our approach integrates several key components: (1) structured context engineering tailored to customer support agents, (2) systematic human-in-the-loop prompt iteration, (3) rigorous LLM judge evaluation with measured inter-rater agreement and GEPA optimization for consistency, and (4) ideation-to-production validation. A central insight is that evaluation-pipeline quality directly determines iteration velocity. We present results from five production deployments spanning distinct domains: card delivery, debt management, credit-limit support, card management, and product explanation. These deployments deliver consistent customer-satisfaction gains while substantially accelerating iteration. In our card-delivery deployment, large-scale A/B testing yields a 37 percentage-point improvement in AI transactional Net Promoter Score and a 29 percentage-point gain in self-service rate over prior agent variants, alongside a strong correlation between offline simulation metrics and online outcomes, demonstrating that eval-driven development reliably predicts production impact. On most use cases, AI satisfaction reaches within a few percentage points of expert human agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。