arXiv:2503.04675cs.CL2025-03NAACL被引 4

用大模型指导策略与检索,让对话满意度预测更准确可解释。

LLM-guided Plan and Retrieval: A Strategic Alignment for Interpretable User Satisfaction Estimation in Dialogue

  • 通过自然语言策略和大模型知识检索,构建可解释的满意度判断框架。
  • 在三个基准上达到当前最优性能,且推理时无需调用大模型。
  • 适合需要透明决策过程的对话系统评估场景。

理解用户对对话系统的满意度(即用户满意度估计,USE)对于评估对话质量与提升用户体验至关重要。然而,现有方法因难以理解用户不满的根本原因,以及标注用户意图成本高昂而面临挑战。为此,我们提出PRAISE(用于可解释满意度估计的策略与检索对齐框架),一个高效可解释的用户满意度预测框架。PRAISE包含三个核心模块:策略规划器生成自然语言标准以分类用户满意度;特征检索器利用大语言模型(LLMs)的知识,从对话语句中检索相关特征;评分分析器评估策略预测并判定满意度。实验结果表明,PRAISE在三个USE基准上达到领先性能。此外,该框架通过将语句与策略有效对齐,提供实例级解释;且推理阶段无需调用大模型,效率更高。

原文摘要 · Abstract (English)

Understanding user satisfaction with conversational systems, known as User Satisfaction Estimation (USE), is essential for assessing dialogue quality and enhancing user experiences. However, existing methods for USE face challenges due to limited understanding of underlying reasons for user dissatisfaction and the high costs of annotating user intentions. To address these challenges, we propose PRAISE (Plan and Retrieval Alignment for Interpretable Satisfaction Estimation), an interpretable framework for effective user satisfaction prediction. PRAISE operates through three key modules. The Strategy Planner develops strategies, which are natural language criteria for classifying user satisfaction. The Feature Retriever then incorporates knowledge on user satisfaction from Large Language Models (LLMs) and retrieves relevance features from utterances. Finally, the Score Analyzer evaluates strategy predictions and classifies user satisfaction. Experimental results demonstrate that PRAISE achieves state-of-the-art performance on three benchmarks for the USE task. Beyond its superior performance, PRAISE offers additional benefits. It enhances interpretability by providing instance-level explanations through effective alignment of utterances with strategies. Moreover, PRAISE operates more efficiently than existing approaches by eliminating the need for LLMs during the inference phase.

对话系统可解释性大模型应用满意度预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。