提出自适应用户模拟器,提升对话推荐系统评估的灵活性与真实性。
Towards Fast Domain Adaptation and Fine-Grained User Simulation for Evaluating Conversational Recommender Systems

- 自动提示生成+开放动作机制,实现跨领域快速适配
- 通过分步思考策略生成多样语言风格,还原用户细微偏好
- 基于广度优先对比的新评估框架,全面检验系统能力
对话推荐系统(CRS)通过多轮交互提升用户体验,但其性能评估仍具挑战。现有基于大语言模型的用户模拟器存在三大缺陷:(1) 领域适应性差,依赖固定提示和预定义动作空间,难以迁移至新领域;(2) 用户建模能力有限,无法准确复现细微语言风格与动态偏好;(3) 评估有效性不足,难以充分检验系统核心能力与鲁棒性。为此,我们提出 AdaptSim——一种自适应领域与自动提示调优的用户模拟器。该方法通过自动提示生成与开放动作机制,降低人工成本,增强跨领域灵活性;在回复生成中采用“先思考后回应”的可控文本生成策略,实现语言风格的细粒度控制;在系统评估方面,引入基于广度优先搜索(BFS)的逐轮成对比较框架,实现全面评估。在三个领域、四种大语言模型上的大量实验表明,AdaptSim 能生成高度真实的对话,显著提升 CRS 能力与鲁棒性评估的有效性与可靠性。
原文摘要 · Abstract (English)
Conversational Recommender Systems (CRSs) enhance user experience through multi-turn interactions, yet evaluating their performance remains challenging. While Large Language Model (LLM) based user simulators are effective, they suffer from three key limitations: (1) Lack of Domain Adaptability: Reliance on fixed prompts and predefined action spaces hinders transfer to novel domains; (2) Limited User Modeling: Inability to accurately replicate subtle linguistic styles and dynamic preferences; (3) Insufficient Evaluation Validity: Existing simulators fail to adequately assess fundamental capabilities and system robustness. To overcome these, we propose AdaptSim, an Adaptive domain and automatic prompt tuning User Simulator. AdaptSim offers an efficient framework for evaluating CRSs by enabling realistic behavior modeling and diverse style generation. It leverages automatic prompt generation and an open action mechanism to reduce manual effort and improve cross-domain flexibility. For response generation, we employ controlled text generation with a "think-then-respond" strategy for fine-grained control over language style. For CRS evaluation, AdaptSim incorporates a novel Breadth-First Search (BFS)-based, turn-level pairwise comparison framework for comprehensive assessment. Extensive experiments across three domains and four LLMs demonstrate that AdaptSim generates realistic dialogues, enabling a highly effective and reliable evaluation of CRS capabilities and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。