arXiv:2608.24314cs.AIcs.ET2026-08

用大模型评估语音助手效果,发现其可靠度依赖具体指标和设置。

Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

论文配图:Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
图 1 · 摘自论文原文
  • 对比人类与GPT-4.1、GPT-5在电信零售对话中的评分表现。
  • 不同评估配置下大模型一致性差异明显,部分指标可靠性不足。
  • 适合自动化评估的指标需人工审核,混合模式更可靠。

大规模评估对话式语音助手需兼顾可观测质量与人类特有的情境判断。本文通过对比人类与GPT-4.1、GPT-5在电信与零售场景语音交互中的评分,考察了基于大模型的评估方法(LLM-as-a-Judge)在对话质量与安全性维度的表现。同一交互在三种配置(p0、p1、p2)下被评分,以检验自动化判断对评估设置的敏感性及结果的泛化能力。除整体一致性外,还分析了各指标相关性、评价者一致性及系统性人-模型分歧,识别出可由自动化可靠判断的对话属性,以及仍依赖上下文理解的领域。此外,语音生成、流式传输及ASR、推理、工具调用阶段的错误传播等流程因素也影响评估效果。结果显示,大模型评估可作为大规模语音助手评估的有效组成部分,但其可靠性取决于具体指标与配置,并非普适。该研究提供了识别适配自动化评估指标的实证框架,支持‘大模型负责规模评估,人类处理高阶判断’的混合评估管道。

原文摘要 · Abstract (English)

Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment typically provided by human evaluators. We investigate LLM-as-a-Judge evaluation by comparing human judgments with GPT-4.1 and GPT-5 on telecom and retail voice-agent conversations, across conversational quality and safety dimensions. The same interac- tions are scored under three evaluation configurations, p0, p1, and p2, to test whether automated judgments are sensitive to the evaluation setup and whether observed patterns generalize across configurations and judge models. Beyond aggregate agreement, we examine metric-level correlations, evaluator consistency, and systematic human-LLM disagreement to identify which conver- sational attributes can be judged reliably by automation and which remain sensitive to interpretation and context. Effective voice-agent evaluation is also shaped by pipeline-level factors such as speech generation, streaming, and error propagation across ASR, reasoning, and tool-calling stages, motivating our focus on comparing how human and LLM judges score the same interactions end to end. Our results show that LLM- based evaluation can serve as an effective component of large- scale voice-agent assessment, but that its reliability is metric- and configuration-dependent rather than uniform. This pro- vides an empirical framework for identifying which metrics suit automated evaluation and supports hybrid pipelines in which LLM judges handle scalable assessment while human evaluators remain engaged for metrics that demand contextual interpretation and higher-confidence judgment.

语音评估大模型评测人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。