用大模型构建用户中心的对话推荐系统评估框架
Evaluating Conversational Recommender Systems via Large Language Models: A User-Centric Framework
- 基于大模型模拟用户真实体验,从12个维度评分
- 多角色辩论机制融合专家意见,生成统一评估分
- 在两个数据集上验证,与人工评价高度一致
对话式推荐系统(CRS)融合推荐与对话任务,其评估极具挑战性。现有方法分别使用规则指标评估推荐与对话管理,但难以反映真实用户体验,也无法得出系统整体表现结论。随着CRS在电商、社交和客服中日益重要,如何用单一指标同时衡量推荐准确性和对话质量,成为阻碍该领域发展的关键问题。本文提出基于大语言模型的用户中心评估框架CoRE,包含两大组件:(1) 大模型作为评估者,总结12项影响用户体验的关键因素并逐项打分;(2) 多智能体辩论机制,由普通用户、领域专家、语言学家和人机交互专家四类角色讨论并综合12个维度,生成整体性能评分。我们在两个基准数据集上评估了四个CRS,结果表明CoRE在多数因素及总体评估上与人工评价高度一致,且整体评分显著优于传统规则指标。
原文摘要 · Abstract (English)
Conversational recommender systems (CRSs) integrate both recommendation and dialogue tasks, making their evaluation uniquely challenging. Existing approaches primarily assess CRS performance by separately evaluating item recommendation and dialogue management using rule-based metrics. However, these methods fail to capture the real human experience, and they cannot draw direct conclusions about the system's overall performance. As conversational recommender systems become increasingly vital in e-commerce, social media, and customer support, the ability to evaluate both recommendation accuracy and dialogue management quality using a single metric, thereby authentically reflecting user experience, has become the principal challenge impeding progress in this field. In this work, we propose a user-centric evaluation framework based on large language models (LLMs) for CRSs, namely Conversational Recommendation Evaluator (CoRE). CoRE consists of two main components: (1) LLM-As-Evaluator. Firstly, we comprehensively summarize 12 key factors influencing user experience in CRSs and directly leverage LLM as an evaluator to assign a score to each factor. (2) Multi-Agent Debater. Secondly, we design a multi-agent debate framework with four distinct roles (common user, domain expert, linguist, and HCI expert) to discuss and synthesize the 12 evaluation factors into a unified overall performance score. Furthermore, we apply the proposed framework to evaluate four CRSs on two benchmark datasets. The experimental results show that CoRE aligns well with human evaluation in most of the 12 factors and the overall assessment. Especially, CoRE's overall evaluation scores demonstrate significantly better alignment with human feedback compared to existing rule-based metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。