用交互式工具帮推荐系统研究者选对数据集,避免实验结果失真。
APS Explorer: Navigating Algorithm Performance Spaces for Informed Dataset Selection
- 通过性能模式可视化数据集相似性,直观展示不同数据间的差异。
- 支持动态对比数据集元特征,辅助选择匹配算法需求的数据集。
- 适合做离线推荐实验的研究者,尤其关注数据与算法匹配性的人。
数据集选择对离线推荐系统实验至关重要,数据不匹配(如稀疏交互场景需低用户-物品密度数据)会导致不可靠结果。然而,86%的ACM RecSys 2024论文未说明数据集选择依据,多数仅使用四个数据集:Amazon(38%)、MovieLens(34%)、Yelp(15%)、Gowalla(12%)。尽管算法性能空间(APS)被提出用于指导数据集选择,但因缺乏直观、交互式工具,其应用受限。为此,我们提出APS Explorer,一个基于网页的交互式可视化工具,支持数据驱动的数据集选择。该工具包含三项功能:(1)交互式PCA图,通过性能模式展现数据集相似性;(2)动态元特征表,支持数据集多维对比;(3)专门的双算法性能对比视图。
原文摘要 · Abstract (English)
Dataset selection is crucial for offline recommender system experiments, as mismatched data (e.g., sparse interaction scenarios require datasets with low user-item density) can lead to unreliable results. Yet, 86\% of ACM RecSys 2024 papers provide no justification for their dataset choices, with most relying on just four datasets: Amazon (38\%), MovieLens (34\%), Yelp (15\%), and Gowalla (12\%). While Algorithm Performance Spaces (APS) were proposed to guide dataset selection, their adoption has been limited due to the absence of an intuitive, interactive tool for APS exploration. Therefore, we introduce the APS Explorer, a web-based visualization tool for interactive APS exploration, enabling data-driven dataset selection. The APS Explorer provides three interactive features: (1) an interactive PCA plot showing dataset similarity via performance patterns, (2) a dynamic meta-feature table for dataset comparisons, and (3) a specialized visualization for pairwise algorithm performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。