arXiv:2603.10570cs.CL2026-03

用自动评估系统降低人工审核成本,提升聊天机器人评价效率。

End-to-End Chatbot Evaluation with Adaptive Reasoning and Uncertainty Filtering

  • 从知识库自动生成问答对,用大模型判断回复正确性
  • 通过置信度过滤不确定答案,与人工判断高度一致
  • 模块化设计支持多领域、多语言快速部署

大语言模型结合检索增强生成已推动领域专用聊天机器人的部署,但系统仍易产生无依据或错误回答。可靠评估至关重要,但人工评审成本高,现有框架常依赖精心构建的测试集和静态指标,难以扩展。本文提出一个端到端自动评估系统,可显著减少人工参与。该系统直接从底层知识库生成问答对,利用大模型对比聊天机器人回复与参考答案,并通过置信度筛选不确定案例。在越南新闻数据集上的应用显示,该评估器与人工判断高度一致,同时大幅降低审查负担。框架具备模块化与语言无关特性,可轻松适配多种领域。本工作提供了一种低人工依赖、可扩展的聊天机器人评估方案。

原文摘要 · Abstract (English)

Large language models (LLMs) combined with retrieval augmented generation have enabled the deployment of domain-specific chatbots, but these systems remain prone to generating unsupported or incorrect answers. Reliable evaluation is therefore critical, yet manual review is costly and existing frameworks often depend on curated test sets and static metrics, limiting scalability. We propose an end-to-end automatic evaluator designed to substantially reduce human effort. Our system generates Q\&A pairs directly from the underlying knowledge base, uses LLMs to judge chatbot responses against reference answers, and applies confidence-based filtering to highlight uncertain cases. Applied to a Vietnamese news dataset, the evaluator achieves high agreement with human judgments while significantly lowering review overhead. The framework is modular and language-agnostic, making it readily adaptable to diverse domains. This work introduces a practical, scalable solution for evaluating chatbots with minimal reliance on manual intervention.

聊天机器人自动评估大模型知识库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。