arXiv:2510.02331cs.CLcs.AI2025-10被引 1

用行为模拟器生成真实用户风格的对话数据,解决推荐系统训练数据不足问题。

Synthetic Dialogue Generation for Interactive Conversational Elicitation & Recommendation (ICER)

  • 通过行为模拟器+大模型提示生成符合用户状态的自然对话
  • 构建大规模开源对话数据集,支持偏好挖掘与评论反馈
  • 人工评估显示对话具高一致性、事实性和自然度,适合推荐系统训练

尽管语言模型在对话推荐系统中潜力巨大,但公开的对话推荐数据稀缺,制约了其微调应用。为此,本文提出一种方法:利用语言模型作为用户模拟器生成数据,结合行为模拟器与提示工程,生成与用户潜在状态一致的自然对话。我们基于此构建了一个大规模、开源的对话推荐数据集,涵盖偏好获取与示例批评两类任务。人工评估表明,部分生成对话在一致性、事实性和自然度方面表现良好,具备实际可用性。

原文摘要 · Abstract (English)

While language models (LMs) offer great potential for conversational recommender systems (CRSs), the paucity of public CRS data makes fine-tuning LMs for CRSs challenging. In response, LMs as user simulators qua data generators can be used to train LM-based CRSs, but often lack behavioral consistency, generating utterance sequences inconsistent with those of any real user. To address this, we develop a methodology for generating natural dialogues that are consistent with a user's underlying state using behavior simulators together with LM-prompting. We illustrate our approach by generating a large, open-source CRS data set with both preference elicitation and example critiquing. Rater evaluation on some of these dialogues shows them to exhibit considerable consistency, factuality and naturalness.

对话生成推荐系统数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。