arXiv:2606.19698cs.CL2026-06

用大模型分析客服对话,比情绪分析更准地判断客户是否真正被解决难题。

What sentiment analysis can't see: Measuring whether customers were helped, and what went wrong, across 70,000 support conversations

  • 用GPT-5.4同时评估语气、满意度和具体问题,超越传统情绪分析。
  • 满意度预测相关系数达0.47,显著高于情绪分析的0.36,误报率更低。
  • 能发现情绪与满意度不一致的44%对话,识别出隐藏的“可修复摩擦”用户。

多数企业通过情绪分析大规模解读客户支持数据,但仅关注客户语气而非实际满意度。本研究在一家领先在线筹款平台的70,450条支持对话上测试了一种更全面的方法:在分析语气的同时,使用GPT-5.4估算客户满意度并标记是否报告了具体问题,并将三种判断结果与客户给出的1至5分评价进行验证。结果显示,满意度预测与真实评分的相关性达到0.47,远高于情绪分析的0.36,且误报率更低。结构化分析揭示了情绪与满意度不一致的情况占44%,单一“中性”标签掩盖了从默默满意到悄然放弃的多种状态;最大群体为“容忍摩擦者”,即虽满意但仍存在可修复问题的客户,此类信息无法通过基于情绪的仪表板发现。研究表明,基于大模型的标注可捕捉语言情感之外的深层信息,为以客户状态和问题根源为核心的新型业务指标提供可能。

原文摘要 · Abstract (English)

Most companies read their customer support data at scale using sentiment analysis, which measures how customers sound rather than whether they were satisfied with the result. We tested a richer alternative on 70,450 support conversations from a leading online fundraising platform: alongside tone, we used GPT-5.4 to estimate each customer's satisfaction and to flag whether they reported a concrete problem, then validated all three readings against the 1-to-5 ratings customers left on the conversations they rated. The satisfaction estimate tracked those ratings far better than sentiment did, correlating at 0.47 against 0.36 and flagging unhappy customers with far fewer false alarms. The structured read also sees what sentiment cannot: tone and satisfaction disagree in 44% of conversations, a single "Neutral" label hides everything from quietly satisfied customers to ones who quietly gave up, and the largest group of all is "tolerated friction," customers who are satisfied but still reporting a fixable problem, a standing issue that no sentiment-based dashboard can surface. The broader finding is that LLM-based annotation can capture far more than the tonality of a customer's language, offering strong potential for new business metrics grounded instead in the customer's state (whether they were satisfied) and the cause of their problem extracted directly from the raw textual data of interactions and feedback.

客户满意度大模型应用对话分析商业智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。