arXiv:2607.12085cs.AI2026-07

构建可扩展的对话系统评估流水线,实现多维度自动质检。

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

  • 配置驱动的流水线,支持异步处理与分片评估
  • 日均处理5万条记录,200万次交互已评估,准确率达89%
  • 适合需高可靠、可审计的零售对话系统研发团队

评估零售对话系统需超越词面匹配,涵盖意图对齐、事实性、帮助性、清晰度、语气一致性及整体质量。尽管大模型评分提供可扩展方案,但生产部署面临治理、复现性、成本、模式一致性、可追溯性和可靠性挑战。本文提出GenAI Evaluation,一个受控的配置驱动流水线,用于大规模评估零售对话系统。该流程对生产聊天日志进行归一化、分片、异步执行,并基于模式约束的大模型评分,评估帮助性、真实性、清晰度、语气一致性及翻译特定维度。仅对不完整、格式错误或模式无效记录进行选择性重评,通过模式锁定、版本化配置、验证日志和记录级溯源保障可审计性。系统每日处理约5万条记录,累计评估超两百万次交互。验证使用4名标注员提供的12,980条分层随机人工标注数据,覆盖14个意图、156个子意图、18个主领域和129个子领域,宏F1达0.93,翻译任务人类接受度达89%。

原文摘要 · Abstract (English)

Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality. Although LLM-as-a-judge methods provide scalable alternatives to human evaluation, production deployment introduces challenges in governance, reproducibility, cost, schema consistency, traceability, and reliability. We present GenAI Evaluation, a governed, configuration-driven pipeline for large-scale evaluation of retail conversational systems. It processes production chatbot logs through normalization, sharding, asynchronous execution, and schema-constrained LLM scoring. The framework evaluates helpfulness, truthfulness, clarity, tone alignment, and translation-specific dimensions. Selective re-evaluation processes only incomplete, malformed, or schema-invalid records, while schema locking, versioned configurations, validation logs, and record-level provenance support auditability. The framework processes approximately 50,000 records daily and has evaluated more than two million interactions. Validation used 12,980 stratified-random human-labeled records from four trained annotators. Classification covered 14 intents, 156 sub-intents, 18 major domains, and 129 sub-domains. The pipeline achieved a macro F1 score of 0.93 and 89% human-acceptability accuracy for translation.

对话评估大模型评分自动化测试零售AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。