构建对话推荐系统真实度评估新基准,解决大模型模拟用户不靠谱的问题。
ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders
- 通过好人/坏人推荐器双轨采集真实对话数据,还原多样用户反馈。
- 发现所有模拟器在真实场景中表现均不佳,存在显著现实差距。
- 提出多维度验证框架,适合评估对话推荐系统的用户模拟能力。
基于大模型的用户模拟器虽有望提升对话式AI性能,但因存在显著“现实差距”,导致系统优化仅针对模拟交互,难以在真实场景中有效运行。本文提出ConvApparel——一个面向人机对话推荐的新数据集,采用独特双代理数据收集机制(同时使用‘优质’和‘劣质’推荐系统),捕捉用户在不同情境下的真实反应,并附带第一人称满意度标注。我们设计了一套综合验证框架,融合统计一致性、人类相似度评分与反事实验证,以测试模型泛化能力。实验表明,所有用户模拟器均存在明显现实差距。然而,数据驱动型模拟器优于提示基线,在反事实验证中更能合理适应未见行为,表明其具备更稳健(尽管不完美)的用户建模能力。
原文摘要 · Abstract (English)
The promise of LLM-based user simulators to improve conversational AI is hindered by a critical "realism gap," leading to systems that are optimized for simulated interactions, but may fail to perform well in the real world. We introduce ConvApparel, a new dataset of human-AI conversations designed to address this gap. Its unique dual-agent data collection protocol -- using both "good" and "bad" recommenders -- enables counterfactual validation by capturing a wide spectrum of user experiences, enriched with first-person annotations of user satisfaction. We propose a comprehensive validation framework that combines statistical alignment, a human-likeness score, and counterfactual validation to test for generalization. Our experiments reveal a significant realism gap across all simulators. However, the framework also shows that data-driven simulators outperform a prompted baseline, particularly in counterfactual validation where they adapt more realistically to unseen behaviors, suggesting they embody more robust, if imperfect, user models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。