arXiv:2602.19339cs.IRcs.LG2026-02

拆解推荐系统数据划分陷阱,让评估更透明可比。

SplitLight: An Exploratory Toolkit for Recommender Systems Datasets and Splits

  • 提供可交互的工具链,分析数据分割中的隐藏问题
  • 检测时间泄露、冷启动暴露等关键错误,提升评估可信度
  • 适合研究者与工程师验证和报告数据预处理方案

推荐系统离线评估常受制于数据准备中未明示的隐含选择。过滤策略、重复处理、冷启动应对及划分方法设计等看似微小的决策,可能显著改变模型排名,破坏复现性与跨论文可比性。本文提出SplitLight——一个开源探索性工具包,使研究者和从业者能够量化、比较并报告预处理与划分流程。给定交互日志及划分子集后,SplitLight分析核心与时间统计特征,刻画重复消费模式与时间戳异常,并诊断划分有效性,包括时间泄露、冷用户/物品暴露及分布偏移。该工具支持多策略并行对比,通过综合汇总与交互可视化呈现结果。以Python库与无代码界面形式交付,生成审计摘要,支撑推荐系统研究与工业应用中的透明、可靠、可比实验。

原文摘要 · Abstract (English)

Offline evaluation of recommender systems is often affected by hidden, under-documented choices in data preparation. Seemingly minor decisions in filtering, handling repeats, cold-start treatment, and splitting strategy design can substantially reorder model rankings and undermine reproducibility and cross-paper comparability. In this paper, we introduce SplitLight, an open-source exploratory toolkit that enables researchers and practitioners designing preprocessing and splitting pipelines or reviewing external artifacts to make these decisions measurable, comparable, and reportable. Given an interaction log and derived split subsets, SplitLight analyzes core and temporal dataset statistics, characterizes repeat consumption patterns and timestamp anomalies, and diagnoses split validity, including temporal leakage, cold-user/item exposure, and distribution shifts. SplitLight further allows side-by-side comparison of alternative splitting strategies through comprehensive aggregated summaries and interactive visualizations. Delivered as both a Python toolkit and an interactive no-code interface, SplitLight produces audit summaries that justify evaluation protocols and support transparent, reliable, and comparable experimentation in recommender systems research and industry.

推荐系统数据分割可复现性工具包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。