提出评估推荐系统稳定与适应性的新方法
Measuring the stability and plasticity of recommender systems
- 通过再训练测试模型对历史模式的保留能力
- 不同算法在稳定性和适应性上表现各异,存在权衡
- 适合关注模型长期表现的系统设计者
传统离线评估仅提供静态性能快照,无法反映在线系统随时间演化的特性。本文提出一种新评估框架,通过再训练机制考察推荐模型在保持历史模式(稳定性)和快速适应变化(可塑性)两方面的表现。该方法不依赖特定数据集、算法或评价指标,具有通用性。在GoodReads数据集上的初步实验表明,三类不同算法展现出显著不同的稳定-可塑性特征,并揭示二者间可能存在权衡。研究还讨论了该框架的潜力与局限,并提出改进方向。
原文摘要 · Abstract (English)
The typical offline protocol to evaluate recommendation algorithms is to collect a dataset of user-item interactions and then use a part of this dataset to train a model, and the remaining data to measure how closely the model recommendations match the observed user interactions. This protocol is straightforward, useful and practical, but it only provides snapshot performance. We know, however, that online systems evolve over time. In general, it is a good idea that models are frequently retrained with recent data. But if this is the case, to what extent can we trust previous evaluations? How will a model perform when a different pattern (re)emerges? In this paper we propose a methodology to study how recommendation models behave when they are retrained. The idea is to profile algorithms according to their ability to, on the one hand, retain past patterns - stability - and, on the other hand, (quickly) adapt to changes - plasticity. We devise an offline evaluation protocol that provides detail on the long-term behavior of models, and that is agnostic to datasets, algorithms and metrics. To illustrate the potential of this framework, we present preliminary results of three different types of algorithms on the GoodReads dataset that suggest different stability and plasticity profiles depending on the algorithmic technique, and a possible trade-off between stability and plasticity. We further discuss the potential and limitations of the proposal and advance some possible improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。