升级用户模拟工具,更好评估对话推荐系统性能
UserSimCRS v2: Simulation-Based Evaluation for Conversational Recommender Systems

- 用基于议程的模拟器+大模型构建更真实用户
- 支持更多推荐系统与数据集,提升评估覆盖范围
- 新增大模型评判工具,适合研究对话推荐的学者
对话推荐系统(CRS)的仿真评估资源匮乏。为此,我们推出了UserSimCRS工具包以填补这一空白。本文介绍UserSimCRS v2,一项重大升级,使其与前沿研究保持一致。主要改进包括:增强的基于议程的用户模拟器、引入基于大语言模型的模拟器、支持更广泛的CRS与数据集,以及新增大模型作为评判者的评估工具。我们在案例研究中展示了这些扩展的有效性。
原文摘要 · Abstract (English)
Resources for simulation-based evaluation of conversational recommender systems (CRSs) are scarce. The UserSimCRS toolkit was introduced to address this gap. In this work, we present UserSimCRS v2, a significant upgrade aligning the toolkit with state-of-the-art research. Key extensions include an enhanced agenda-based user simulator, introduction of large language model-based simulators, integration for a wider range of CRSs and datasets, and new LLM-as-a-judge evaluation utilities. We demonstrate these extensions in a case study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。