arXiv:2508.00889cs.CLcs.AI2025-08中稿 · KDD

针对客服对话分析中AI生成结论的幻觉问题,提出可落地的评估框架与数据集。

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts

  • 设计'分解-解耦-剥离'三步评估法,让事实性判断有语言学依据。
  • 构建首个面向客服对话解读的真相性评测数据集FECT,含人工标注标准。
  • 验证大模型在新框架下可对解读类输出进行可靠事实性判断,适合企业级应用。

大型语言模型(LLMs)常产生与输入、参考材料或现实知识不符的幻觉内容。在企业决策支持场景中,此类错误尤为严重。用于分析客服对话并生成摘要的LLM面临独特挑战:关于情感倾向和业务问题根因的分析结论往往缺乏真实标签。为此,我们首次提出3D范式——分解、解耦、剥离,应用于人工标注指南与大模型评判提示,使事实性标签建立在语言学基础上。随后构建了全新的基准数据集FECT,用于评估在客服对话转录文本中生成的解释性陈述的事实性。最后报告了将大模型对齐至3D范式的实验结果。整体研究为自动评估客服对话分析系统输出的真实性提供了新方法。

原文摘要 · Abstract (English)

Large language models (LLMs) are known to hallucinate, producing natural language outputs that are not grounded in the input, reference materials, or real-world knowledge. In enterprise applications where AI features support business decisions, such hallucinations can be particularly detrimental. LLMs that analyze and summarize contact center conversations introduce a unique set of challenges for factuality evaluation, because ground-truth labels often do not exist for analytical interpretations about sentiments captured in the conversation and root causes of the business problems. To remedy this, we first introduce a \textbf{3D} -- \textbf{Decompose, Decouple, Detach} -- paradigm in the human annotation guideline and the LLM-judges' prompt to ground the factuality labels in linguistically-informed evaluation criteria. We then introduce \textbf{FECT}, a novel benchmark dataset for \textbf{F}actuality \textbf{E}valuation of Interpretive AI-Generated \textbf{C}laims in Contact Center Conversation \textbf{T}ranscripts, labeled under our 3D paradigm. Lastly, we report our findings from aligning LLM-judges on the 3D paradigm. Overall, our findings contribute a new approach for automatically evaluating the factuality of outputs generated by an AI system for analyzing contact center conversations.

事实性评估客服对话大模型幻觉数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。