用大模型拆解9000次客服对话,发现满意度维度间高度相关但产品项不准确。
Dimensionality in Satisfaction Ratings
- 用GPT-4.1将满意度分解为五维:整体、客服、结果、产品、客户付出
- 除产品外四维与自评相关性达0.65,剔除极端分歧后相关性升至0.914
- 全量分析显示满意度比抽样调查低20%,揭示传统数据偏差
我们使用大型语言模型(GPT-4.1)对一家全球消费品公司的约9000条客服对话文本进行标注,将客户满意度分解为五个维度:整体、客服人员、结果、产品及客户付出。通过与客户自评满意度对比验证了标注有效性。其中四个维度与自评满意度高度一致(整体、客服、结果相关系数接近0.65,客户付出为-0.54),而产品满意度与现有代理指标关联较弱。未调整的相关性低估了实际一致性:分歧主要集中在少数可读的异常对话中,若排除严重分歧,总体相关性提升至0.811;若排除全部异常尾部,可达0.914。各维度之间高度共线,将它们加入整体评分并未提升对客户评分的预测能力。因此,该分解的价值在于归因分析和覆盖范围扩展,而非增量预测。更重要的是,当分析所有通话记录而非仅回复问卷的少数样本时,客户满意度显著低于问卷报告值(全量普查得分为2.91,抽样调查为3.62,均在五分制上)。该方法的潜力在于从对话数据中识别更精细的客户体验驱动因素。
原文摘要 · Abstract (English)
We used a large language model (GPT-4.1) to annotate the text of about 9,000 support conversations at a global consumer-goods firm, decomposing customer-care satisfaction into component axes (overall, agent, outcome, product, and customer effort), and validated the LLM annotations against the satisfaction ratings customers gave themselves. Four of five axes track self-reported satisfaction closely (overall, agent, and outcome near an unadjusted 0.65; effort -0.54), while product satisfaction is weak against the available proxy. The unadjusted correlation also understates the alignment: the disagreements concentrate in a small, readable tail of divergent sessions rather than in general drift, and the overall correlation rises to 0.811 when only the severe divergences are excluded and to 0.914 when the full divergent tail is excluded. The axes are also highly collinear, and adding them to the overall score does not improve prediction of the customer's rating, the decomposition's value is not incremental prediction but attribution and coverage. And, with greater coverage the picture of the data changes. Read on every contact rather than the few that return a survey, satisfaction is markedly lower than the survey reports (a full-census 2.91 against the surveyed 3.62 on a five-point scale). The promise of decomposed satisfaction as a methodology is the ability to identify more nuanced drivers of customer experience in conversational data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。