用模型不确定性估算推荐效果,无需真实标签。
Are Recommenders Self-Aware? Label-Free Recommendation Performance Estimation via Model Uncertainty
- 基于物品预测分布计算推荐列表的概率分布,衡量模型不确定性。
- 在合成与真实数据上,其相关性优于其他无标签评估方法。
- 可观察模型训练与推理中的状态变化,适合构建透明推荐系统。
推荐模型能否自我感知?本文通过量化模型不确定性,实现无需标签的性能估计,使推荐系统在面向用户前即可自评估。提出概率型列表分布不确定性(LiDu)方法,通过个体物品预测分布推导生成特定推荐列表的概率。在矩阵分解模型与真实数据集上的实验表明,LiDu相较于多种无标签评估方法,与实际推荐性能的相关性更高。此外,该方法能揭示模型在训练与推理过程中的动态状态变化。本工作建立了推荐不确定性与性能之间的实证关联,为构建更透明、可自评的推荐系统迈出关键一步。
原文摘要 · Abstract (English)
Can a recommendation model be self-aware? This paper investigates the recommender's self-awareness by quantifying its uncertainty, which provides a label-free estimation of its performance. Such self-assessment can enable more informed understanding and decision-making before the recommender engages with any users. To this end, we propose an intuitive and effective method, probability-based List Distribution uncertainty (LiDu). LiDu measures uncertainty by determining the probability that a recommender will generate a certain ranking list based on the prediction distributions of individual items. We validate LiDu's ability to represent model self-awareness in two settings: (1) with a matrix factorization model on a synthetic dataset, and (2) with popular recommendation algorithms on real-world datasets. Experimental results show that LiDu is more correlated with recommendation performance than a series of label-free performance estimators. Additionally, LiDu provides valuable insights into the dynamic inner states of models throughout training and inference. This work establishes an empirical connection between recommendation uncertainty and performance, framing it as a step towards more transparent and self-evaluating recommender systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。