arXiv:2510.05624cs.IR2025-10被引 9

现有对话推荐系统评估方法存在缺陷,用户模拟可提升评估真实性。

Limitations of Current Evaluation Practices for Conversational Recommender Systems and the Potential of User Simulation

  • 用用户模拟生成动态交互数据,替代静态测试集。
  • 新指标与真实用户满意度相关性更高,验证了有效性。
  • 适合研究评估方法、系统设计的学者和开发者参考。

对话推荐系统(CRS)的研究与发展依赖于可靠的评估方法,但其交互特性给自动化评估带来挑战。本文批判性分析现有评估实践,指出两大局限:过度依赖静态测试集,以及现有评估指标不足。通过分析九个现有CRS的真实用户交互,我们发现用户自评满意度与以往文献中的性能分数存在显著脱节。为此,本文探索用户模拟技术生成动态交互数据,突破静态数据限制,并提出基于通用奖励/成本框架的新评估指标,更贴近真实用户满意度。对不同模拟方法的分析表明,其能有效提升评估结果与人工评价排名的相关性,初步验证了可行性。尽管如此,模拟技术和评估指标仍有待进一步研究和完善。

原文摘要 · Abstract (English)

Research and development on conversational recommender systems (CRSs) critically depends on sound and reliable evaluation methodologies. However, the interactive nature of these systems poses significant challenges for automatic evaluation. This paper critically examines current evaluation practices and identifies two key limitations: the over-reliance on static test collections and the inadequacy of existing evaluation metrics. To substantiate this critique, we analyze real user interactions with nine existing CRSs and demonstrate a striking disconnect between self-reported user satisfaction and performance scores reported in prior literature. To address these limitations, this work explores the potential of user simulation to generate dynamic interaction data, offering a departure from static datasets. Furthermore, we propose novel evaluation metrics, based on a general reward/cost framework, designed to better align with real user satisfaction. Our analysis of different simulation approaches provides valuable insights into their effectiveness and reveals promising initial results, showing improved correlation with system rankings compared to human evaluation. While these findings indicate a significant step forward in CRS evaluation, we also identify areas for future research and refinement in both simulation techniques and evaluation metrics.

对话推荐评估方法用户模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。