系统梳理对话推荐系统的用户体验研究,揭示评估盲区与LLM带来的新挑战。
Evaluating User Experience in Conversational Recommender Systems: A Systematic Review Across Classical and LLM-Powered Approaches
- 基于PRISMA指南分析23项实证研究,涵盖经典与大模型驱动的推荐系统。
- 发现后置问卷主导评估,对话层级情感体验和自适应行为与体验关联薄弱。
- 提出面向大模型的用户体验评估框架,适合人机交互与推荐系统研究者。
对话推荐系统(CRSs)在多个领域受到越来越多关注,但其用户体验(UX)评估仍显不足。现有综述普遍忽略实证性用户体验研究,尤其是自适应及大语言模型(LLM)驱动的系统。为填补这一空白,我们遵循PRISMA指南开展系统性综述,整合了2017至2025年间发表的23项实证研究。分析了用户体验的概念化、测量方式及其受领域、自适应性及LLM影响的程度。研究发现:后置问卷占主导地位,对话层级的情感体验指标很少被评估,自适应行为与用户体验结果之间的关联也极少被探讨。大模型驱动的系统引入了认知不透明性和冗长输出等新问题,但评估中对此类问题关注不足。本文贡献包括一套结构化的用户体验度量体系、对自适应与非自适应系统间的比较分析,以及面向未来的大模型感知型用户体验评估议程。研究成果有助于推动更透明、更具参与感且以用户为中心的对话推荐系统评估实践。
原文摘要 · Abstract (English)
Conversational Recommender Systems (CRSs) are receiving growing research attention across domains, yet their user experience (UX) evaluation remains limited. Existing reviews largely overlook empirical UX studies, particularly in adaptive and large language model (LLM)-based CRSs. To address this gap, we conducted a systematic review following PRISMA guidelines, synthesising 23 empirical studies published between 2017 and 2025. We analysed how UX has been conceptualised, measured, and shaped by domain, adaptivity, and LLM. Our findings reveal persistent limitations: post hoc surveys dominate, turn-level affective UX constructs are rarely assessed, and adaptive behaviours are seldom linked to UX outcomes. LLM-based CRSs introduce further challenges, including epistemic opacity and verbosity, yet evaluations infrequently address these issues. We contribute a structured synthesis of UX metrics, a comparative analysis of adaptive and nonadaptive systems, and a forward-looking agenda for LLM-aware UX evaluation. These findings support the development of more transparent, engaging, and user-centred CRS evaluation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。