arXiv:2604.00022cs.CLcs.AI2026-04被引 1

用大模型评分对话质量,发现不同维度对转化率影响差异大。

Criterion Validity of LLM-as-Judge for Business Outcomes in Conversational Commerce

论文配图:Criterion Validity of LLM-as-Judge for Business Outcomes in Conversational Commerce
图 1 · 摘自论文原文
  • 用7维评分体系评估对话,通过大模型打分
  • 需求挖掘和节奏策略与转化显著相关,记忆维度无关
  • 评分权重应按实际效果调整,避免整体评分被稀释

多维度评分体系广泛用于对话AI评估,但其准则效度——即评分是否真实反映下游业务结果——尚未验证。本研究基于中国某主流婚恋平台开展两阶段研究,测试由大模型作为裁判的7维评分体系与真实转化之间的关联。在第二阶段(n=60条人工对话,分层抽样,有验证标签)中,需求挖掘(D1:rho=0.368,p=0.004)与节奏策略(D3:rho=0.354,p=0.006)显著正相关于转化,而上下文记忆(D5:rho=0.018,不显著)无明显关联。等权综合评分(rho=0.272)低于最优单项,存在综合稀释效应;基于转化优化权重后,评分相关性提升至0.351。控制对话长度后,D3的关联仍显著增强(OR=3.18,p=0.006),排除时长干扰。初步试点(n=14)出现的“评估-结果悖论”实为代理类型混淆所致。行为分析(130条对话)结合信任漏斗框架指出:AI代理虽执行销售行为,却未能建立用户信任。研究提出三层评估架构,并倡导将准则效度测试纳入对话评估常规流程。

原文摘要 · Abstract (English)

Multi-dimensional rubric-based dialogue evaluation is widely used to assess conversational AI, yet its criterion validity -- whether quality scores are associated with the downstream outcomes they are meant to serve -- remains largely untested. We address this gap through a two-phase study on a major Chinese matchmaking platform, testing a 7-dimension evaluation rubric (implemented via LLM-as-Judge) against verified business conversion. Our findings concern rubric design and weighting, not LLM scoring accuracy: any judge using the same rubric would face the same structural issue. The core finding is dimension-level heterogeneity: in Phase 2 (n=60 human conversations, stratified sample, verified labels), Need Elicitation (D1: rho=0.368, p=0.004) and Pacing Strategy (D3: rho=0.354, p=0.006) are significantly associated with conversion after Bonferroni correction, while Contextual Memory (D5: rho=0.018, n.s.) shows no detectable association. This heterogeneity causes the equal-weighted composite (rho=0.272) to underperform its best dimensions -- a composite dilution effect that conversion-informed reweighting partially corrects (rho=0.351). Logistic regression controlling for conversation length confirms D3's association strengthens (OR=3.18, p=0.006), ruling out a length confound. An initial pilot (n=14) mixing human and AI conversations had produced a misleading "evaluation-outcome paradox," which Phase 2 revealed as an agent-type confound artifact. Behavioral analysis of 130 conversations through a Trust-Funnel framework identifies a candidate mechanism: AI agents execute sales behaviors without building user trust. We operationalize these findings in a three-layer evaluation architecture and advocate criterion validity testing as standard practice in applied dialogue evaluation.

对话评估大模型评分转化率效度检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。