arXiv:2607.26060cs.CLcs.AI2026-07

用数字孪生客户模拟大模型客服,实现低成本高可靠验证

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

论文配图:Large-Scale ChatBot Validation Through Customer Digital Twin Simulations
图 1 · 摘自论文原文
  • 基于真实数据构建高保真虚拟客户,可自动生成并控制行为风格
  • 在多种情绪、人群和语言场景下验证通过,幻觉率低且性格还原准确
  • 适合金融等强监管领域,助力银行合规部署大模型客服

基于大语言模型的聊天机器人正在重塑银行业等受监管领域的客户服务,但规模化、低成本的验证仍是安全部署的关键瓶颈。本文提出两项贡献:一是构建基于真实交易与对话数据的高保真合成客户代理(SCA)数字孪生方法,支持自动创建和行为调控,可模拟多样客户画像与互动风格;评估显示,SCA在语义对齐、幻觉率和人格特征还原方面表现优异,且可通过可控干预实现精准调节。二是开发基于SCA的验证框架,融合自动化大模型评分、专家人工测试与对抗性探查。在情绪状态、人口群体与语言因素等多维度场景下验证均表现稳健。该方法已在英国一家领先银行的实际客户对客聊天机器人验证中应用,为金融机构提供了可扩展的合规路径。

原文摘要 · Abstract (English)

LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. We present a two-part contribution for large-scale chatbot validation. First, we introduce a methodology for creating high-fidelity synthetic customer agents (SCAs) as digital twins, grounded in real transactional and conversational data, that enables automatic generation and behavioral conditioning to simulate diverse customer profiles and interaction styles. Evaluation demonstrates that SCAs achieve high semantic alignment with real customers, low hallucination rates, and successful personality trait reproduction with controllable interventions. Second, we develop an SCA-based validation framework combining automated LLM-as-a-Judge evaluation, human expert testing, and adversarial probing. Scenario-based validation across emotional states, demographic groups, and linguistic factors confirms robust performance. Our approach was used to validate a customer facing chatbot at a leading UK bank, providing financial institutions with a scalable pathway toward regulatory compliance.

大模型验证数字孪生金融AI聊天机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。