arXiv:2607.07985cs.CLcs.AI2026-07被引 2

用AI模型评估语音助手对话质量,效果接近真人且成本低

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

论文配图:A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents
图 1 · 摘自论文原文
  • 用Gemini模型直接分析原始音频波形评分
  • 7项指标与真人评分相关性高,6项一致率超60%
  • 适合大规模语音系统评估,可替代或补充人工评测

我们实证评估了Gemini系列模型(2.5 Flash、3.5 Flash、3.1 Pro)作为音频评判者在全双工对话中的可靠性。以Gemini 2.5 Flash为基准,对比三位校准真人评分员对209段立体声会话的评分,涵盖8个生产维度:152次全双工对话,覆盖13种口音与场景组合,以及57段注入缺陷的片段。结果表明:(i)在8项中的5项,模型与真人间的斯皮尔曼等级相关系数差异不超过0.07,7项的95%置信区间重叠;(ii)在8项中的6项,模型与三名真人平均分的差异在1分以内,覆盖60%至92%的会话;(iii)48个(缺陷×维度)组合中,45个模型敏感度不亚于或优于真人(基于Newcombe-Wilson置信区间),但多数为统计效力不足的零假设。跨模型表现显示:3.5 Flash在全部8项上达成良好一致性,而3.1 Pro虽相关性相近,但在部分维度评分显著低于真人。模型切换需重新校准,不可仅凭相关性推断。识别出四个需谨慎部署的领域。估算显示,当前评估频率下,纯人工评分成本约为LALM工作的百倍。本研究为在支持维度上以LALM替代或补充人工评分提供了实证基础。

原文摘要 · Abstract (English)

We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1 Pro. Our primary evidence base uses Gemini 2.5 Flash as the ground-truth model, validated against three calibrated human raters on 209 stereo sessions, scored on 8 production dimensions: 152 full-duplex conversations across 13 accent-and-condition strata, together with 57 adversarial defect-injected clips. The evidence for Gemini 2.5 Flash is consistent across three tests. (i) On 5 of 8 dimensions the LALM-human Spearman rho departs from the pairwise human-human rho by at most 0.07, and on 7 of 8 dimensions the two quantities 95 percent bootstrap confidence intervals overlap. (ii) The LALM agrees with the three-rater human mean within 1 point on 60 to 92 percent of sessions on 6 of 8 dimensions. (iii) On 45 of 48 (defect, dimension) cells the LALM is as sensitive as humans or better under Newcombe-Wilson 95 percent confidence intervals, though most of these are underpowered nulls rather than demonstrated parity. Rank-ordering ability transfers across the Gemini family: 3.5 Flash improves simple agreement to 8 of 8 dimensions, while 3.1 Pro rates several dimensions markedly lower than humans despite comparable rank correlation. A model swap should be re-validated on calibration specifically, not assumed from rank-correlation alone. We identify four areas where deployment requires care, and we estimate that human rating alone for our current evaluation cadence costs roughly two orders of magnitude more than the equivalent LALM workload. The data presented here provides a defensible empirical basis for deploying the LALM as a substitute or fourth rater on the dimensions where the evidence supports it.

语音评估AI评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。