用大模型评分对话质量,发现不同维度对转化率影响差异大。
Criterion Validity of LLM-as-Judge for Business Outcomes in Conversational Commerce

- 用7维评分体系评估对话,通过大模型打分
- 需求挖掘和节奏策略与转化显著相关,记忆维度无关
- 评分权重应按实际效果调整,避免整体评分被稀释
多维度评分体系广泛用于对话AI评估,但其准则效度——即评分是否真实反映下游业务结果——尚未验证。本研究基于中国某主流婚恋平台开展两阶段研究,测试由大模型作为裁判的7维评分体系与真实转化之间的关联。在第二阶段(n=60条人工对话,分层抽样,有验证标签)中,需求挖掘(D1:rho=0.368,p=0.004)与节奏策略(D3:rho=0.354,p=0.006)显著正相关于转化,而上下文记忆(D5:rho=0.018,不显著)无明显关联。等权综合评分(rho=0.272)低于最优单项,存在综合稀释效应;基于转化优化权重后,评分相关性提升至0.351。控制对话长度后,D3的关联仍显著增强(OR=3.18,p=0.006),排除时长干扰。初步试点(n=14)出现的“评估-结果悖论”实为代理类型混淆所致。行为分析(130条对话)结合信任漏斗框架指出:AI代理虽执行销售行为,却未能建立用户信任。研究提出三层评估架构,并倡导将准则效度测试纳入对话评估常规流程。
原文摘要 · Abstract (English)
Multi-dimensional rubric-based dialogue evaluation is widely used to assess conversational AI, yet its criterion validity -- whether quality scores are associated with the downstream outcomes they are meant to serve -- remains largely untested. We address this gap through a two-phase study on a major Chinese matchmaking platform, testing a 7-dimension evaluation rubric (implemented via LLM-as-Judge) against verified business conversion. Our findings concern rubric design and weighting, not LLM scoring accuracy: any judge using the same rubric would face the same structural issue. The core finding is dimension-level heterogeneity: in Phase 2 (n=60 human conversations, stratified sample, verified labels), Need Elicitation (D1: rho=0.368, p=0.004) and Pacing Strategy (D3: rho=0.354, p=0.006) are significantly associated with conversion after Bonferroni correction, while Contextual Memory (D5: rho=0.018, n.s.) shows no detectable association. This heterogeneity causes the equal-weighted composite (rho=0.272) to underperform its best dimensions -- a composite dilution effect that conversion-informed reweighting partially corrects (rho=0.351). Logistic regression controlling for conversation length confirms D3's association strengthens (OR=3.18, p=0.006), ruling out a length confound. An initial pilot (n=14) mixing human and AI conversations had produced a misleading "evaluation-outcome paradox," which Phase 2 revealed as an agent-type confound artifact. Behavioral analysis of 130 conversations through a Trust-Funnel framework identifies a candidate mechanism: AI agents execute sales behaviors without building user trust. We operationalize these findings in a three-layer evaluation architecture and advocate criterion validity testing as standard practice in applied dialogue evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。