针对对话系统中少数用户满意度估计偏差问题,提出自适应强化学习框架。
Minority-Aware Satisfaction Estimation in Dialogue Systems via Preference-Adaptive Reinforcement Learning
- 用可解释推理链捕捉个体偏好,建模用户独特需求。
- 无监督聚类发现不同用户群体,学习群体级满意度模式。
- 联合优化个体与群体偏好,在少数群体上提升估计准确率。
对话系统中的用户满意度具有主观性。当采用统一响应策略时,由于个体意图和偏好的差异,少数用户可能给出与多数用户不同的满意度评分。现有对齐方法通常训练通用模型以达成广泛共识,常忽略少数群体视角和个性化适配。本文提出统一框架,同时建模个体与群体层面的满意度偏好。首先引入可解释推理链(CoPeR)捕捉个体偏好;其次提出基于期望最大化算法的多数-少数偏好感知聚类(M2PC),无监督发现用户群体并学习群体偏好;最后将两者整合进偏好自适应强化学习框架(PAda-PPO),联合优化个体与群体偏好对齐。在情感支持对话数据集上的实验表明,该方法在用户满意度估计上持续提升,尤其在代表性不足的用户群体中表现更优。
原文摘要 · Abstract (English)
User satisfaction in dialogue systems is inherently subjective. When the same response strategy is applied across users, minority users may assign different satisfaction ratings than majority users due to variations in individual intents and preferences. However, existing alignment methods typically train one-size-fits-all models that aim for broad consensus, often overlooking minority perspectives and user-specific adaptation. We propose a unified framework that models both individual- and group-level preferences for user satisfaction estimation. First, we introduce Chain-of-Personalized-Reasoning (CoPeR) to capture individual preferences through interpretable reasoning chains. Second, we propose an expectation-maximization-based Majority-Minority Preference-Aware Clustering (M2PC) algorithm that discovers distinct user groups in an unsupervised manner to learn group-level preferences. Finally, we integrate these components into a preference-adaptive reinforcement learning framework (PAda-PPO) that jointly optimizes alignment with both individual and group preferences. Experiments on the Emotional Support Conversation dataset demonstrate consistent improvements in user satisfaction estimation, particularly for underrepresented user groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。