对比离线与在线评估,验证推荐系统在伊朗真实场景下的效果
Online and Offline Evaluations of Collaborative Filtering and Content Based Recommender Systems
- 结合离线指标与线上A/B测试,多维度评估推荐算法
- 系统处理300请求/秒,覆盖70个波斯语网站,运行超一年
- 针对冷启动与热门偏差提出改进方法,适合本地化推荐场景
推荐系统是广泛使用的AI应用,旨在帮助用户高效发现相关项目。其有效性与用户及平台满意度密切相关,但用户满意度难以用信息检索和准确率等数学指标完整描述。尽管许多研究通过离线测试评估准确性,越来越多学者认为在线评估(如A/B测试)更合适。我们在不同规模和主题的数据集上应用多种算法,在媒体流媒体、数字出版、电商及新闻广播等多个平台生成推荐。特别地,目标网站与数据集为波斯语(法尔西语)。本研究对一个已持续运行一年、覆盖约70个伊朗网站、总处理量达每秒约300请求的大规模推荐系统进行对比分析。系统采用基于用户与物品的协同过滤、内容基础、趋势驱动及混合方法。通过离线(准确率、命中率@k、nDCG)与在线(点击率CTR)双重评估,识别各算法在特定数据与系统规模下的最优表现,并提出缓解冷启动与流行度偏差的方法。
原文摘要 · Abstract (English)
Recommender systems are widely used AI applications designed to help users efficiently discover relevant items. The effectiveness of such systems is tied to the satisfaction of both users and providers. However, user satisfaction is complex and cannot be easily framed mathematically using information retrieval and accuracy metrics. While many studies evaluate accuracy through offline tests, a growing number of researchers argue that online evaluation methods such as A/B testing are better suited for this purpose. We have employed a variety of algorithms on different types of datasets divergent in size and subject, producing recommendations in various platforms, including media streaming services, digital publishing websites, e-commerce systems, and news broadcasting networks. Notably, our target websites and datasets are in Persian (Farsi) language. This study provides a comparative analysis of a large-scale recommender system that has been operating for the past year across about 70 websites in Iran, processing roughly 300 requests per second collectively. The system employs user-based and item-based recommendations using content-based, collaborative filtering, trend-based methods, and hybrid approaches. Through both offline and online evaluations, we aim to identify where these algorithms perform most efficiently and determine the best method for our specific needs, considering the dataset and system scale. Our methods of evaluation include manual evaluation, offline tests including accuracy and ranking metrics like hit-rate@k and nDCG, and online tests consisting of click-through rate (CTR). Additionally we analyzed and proposed methods to address cold-start and popularity bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。