arXiv:2504.15476cs.IR2025-04被引 7

用非对话数据生成对话推荐系统训练样本,无需真实对话语料

An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems

  • 从评论、元数据等非对话信号中合成对话数据,实现零数据启动
  • 合成数据在低资源场景下表现优于少量真实对话,且提升推荐效果
  • 主动选择优质信号可显著提高数据效率,适合新领域快速部署

对话式推荐系统(CRS)通常依赖特定领域的对话数据,而这类数据成本高、稀缺且在新领域常不可得。本文系统研究了零数据下构建CRS的自举方法:仅利用非对话信号(如商品评论、元数据、用户-物品交互)生成合成对话监督信号,无需任何领域内对话语料。我们对比了两种信息论选择策略(Jensen-Shannon多样性与Fisher信息),涵盖不同领域信号、模型架构、数据集和微调范式。结果表明,基于领域知识的合成数据始终优于零样本提示和简单合成基线;主动选择显著提升数据效率;元数据与协同过滤信号各自改善选择质量;在低资源情况下,合成数据甚至超越稀缺的真实对话,并能进一步补充其效果。这些发现确立了非对话领域信号作为无需对话数据构建CRS的可行路径。代码已公开于 https://anonymous.4open.science/r/zero_data_crs/。

原文摘要 · Abstract (English)

Conversational Recommender Systems (CRS) typically require domain-specific dialogue data, which is costly, scarce, and often unavailable in new domains. We conduct a systematic empirical study of zero-data CRS bootstrapping: generating synthetic conversational supervision from non-conversational signals---item reviews, metadata, and user-item interactions---without any in-domain dialogue corpus. We compare two information-theoretic selection strategies, Jensen-Shannon diversity and Fisher information, across domain signals, model architectures, datasets, and fine-tuning paradigms. Our results show that domain-grounded synthetic data consistently outperforms zero-shot prompting and naive synthetic baselines; active selection improves data efficiency over random sampling; metadata and collaborative filtering signals each improve selection quality; and, in low-resource settings, synthetic data can outperform scarce real dialogues while further complementing them. These findings establish non-conversational domain signals as a viable path toward building CRS without conversational training data. The code is available at https://anonymous.4open.science/r/zero_data_crs/ .

对话推荐零样本数据合成低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。