arXiv:2605.13053cs.IR2026-05中稿 · Proceedings of the…

重评对话推荐系统,发现多数结果依赖重复套路而非真实推荐能力。

A Standardized Re-evaluation of Conversational Recommender Systems on the ReDial Dataset

论文配图:A Standardized Re-evaluation of Conversational Recommender Systems on the ReDial Dataset
图 1 · 摘自论文原文
  • 统一实验条件重测7个主流模型,消除数据预处理差异影响。
  • 近半数准确率来自可复现的重复行为,非真正推荐效果。
  • 大模型能力远超架构创新,建议关注用户交互效率与新颖性。

近年来对话推荐系统(CRS)研究蓬勃发展,其中ReDial数据集被广泛使用,但因预处理方式和真实物品定义不一,导致不同研究间结果难以比较。此外,大语言模型(LLM)选择和外部数据源引入进一步加剧了混淆因素。本文在标准化条件下重新评估三个架构族的七种代表性方法。复现研究表明存在‘粒度差距’:细粒度排序(Recall@1)对实现细节极为敏感;可复现性分析显示,近50%的报告准确率源于‘重复捷径’,在注重新颖性的评估中并不存在。此外,性能提升多源于LLM主干模型容量,而非特定架构创新。通过引入以用户为中心的效用指标,我们证明传统召回率常高估系统实际对话效能。本工作建立透明可控基线,倡导更重视新颖性与交互效率的评估实践。

原文摘要 · Abstract (English)

Recent years have seen a surge of research into conversational recommender systems (CRS). Among existing datasets, ReDial is the most widely used benchmark, cited in hundreds of studies. However, variations in how the dataset is preprocessed and used in experiments, particularly in the definition of ground-truth items, make it difficult to compare results across studies. These comparisons are further complicated by confounding factors such as the choice of the underlying large language model (LLM) and the use of external data sources. In this work, we revisit seven prominent CRS methods across three architectural families and evaluate them under standardized conditions. Our reproducibility study reveals a ``granularity gap,'' where fine-grained ranking (Recall@1) is highly sensitive to implementation details, while our replicability analysis shows that nearly 50% of reported accuracy stems from ``repetition shortcuts'' that are absent in novelty-focused evaluation. Furthermore, we find that performance gains are often driven more by the capacity of the LLM backbone than by specific architectural innovations. Finally, by applying user-centric utility metrics, we demonstrate that traditional recall frequently overstates a system's actual conversational effectiveness. This work establishes a transparent, controlled baseline and promotes evaluation practices that prioritize novelty and interaction efficiency.

对话推荐可复现性评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。