arXiv:2411.16160cs.IR2024-11中稿 · EMNLP被引 6

用真实用户对话数据构建无目标模拟器,让推荐系统真正“问出”用户喜好。

Stop Playing the Guessing Game! Target-free User Simulation for Evaluating Conversational Recommender Systems

  • 基于真实用户交互历史构建无预设目标的对话模拟器
  • 通过多轮互动评估系统发现用户偏好的能力,而非简单猜中目标项
  • 提出四维度综合评价体系,更全面衡量推荐系统对话能力

对话式推荐系统(CRS)近年尝试通过模拟真实用户与系统对话来构建更真实的测试环境。然而,现有评估方法常依赖有预设目标的用户模拟器,导致对话变成简单的猜谜游戏,无法反映真实用户逐步发现偏好的过程。这些模拟器通常依据固定属性引导系统指向特定目标项,限制了动态探索能力,且评估指标多聚焦单轮召回率,忽视偏好获取的中间过程。为此,本文提出PEPPER协议,采用真实用户交互历史和评论构建无目标用户模拟器,使对话自然展开,支持用户在丰富互动中逐步明确偏好,从而更真实地评估系统偏好获取能力。此外,PEPPER引入涵盖定量与定性维度的综合评价指标,覆盖偏好获取过程的四个关键方面。大量实验验证了其有效性,并对现有CRS在偏好获取与推荐方面的表现进行了深入分析。

原文摘要 · Abstract (English)

Recent approaches in Conversational Recommender Systems (CRSs) have tried to simulate real-world users engaging in conversations with CRSs to create more realistic testing environments that reflect the complexity of human-agent dialogue. Despite the significant advancements, reliably evaluating the capability of CRSs to elicit user preferences still faces a significant challenge. Existing evaluation metrics often rely on target-biased user simulators that assume users have predefined preferences, leading to interactions that devolve into simplistic guessing game. These simulators typically guide the CRS toward specific target items based on fixed attributes, limiting the dynamic exploration of user preferences and struggling to capture the evolving nature of real-user interactions. Additionally, current evaluation metrics are predominantly focused on single-turn recall of target items, neglecting the intermediate processes of preference elicitation. To address this, we introduce PEPPER, a novel CRS evaluation protocol with target-free user simulators constructed from real-user interaction histories and reviews. PEPPER enables realistic user-CRS dialogues without falling into simplistic guessing games, allowing users to gradually discover their preferences through enriched interactions, thereby providing a more accurate and reliable assessment of the CRS's ability to elicit personal preferences. Furthermore, PEPPER presents detailed measures for comprehensively evaluating the preference elicitation capabilities of CRSs, encompassing both quantitative and qualitative measures that capture four distinct aspects of the preference elicitation process. Through extensive experiments, we demonstrate the validity of PEPPER as a simulation environment and conduct a thorough analysis of how effectively existing CRSs perform in preference elicitation and recommendation.

对话推荐用户模拟偏好获取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。