arXiv:2412.06809cs.IRcs.AI2024-12

生成高维稀疏的多样化合成数据集,用于真实推荐系统评估。

Generating Diverse Synthetic Datasets for Evaluation of Real-life Recommender Systems

  • 构建可调控属性的模块化合成数据生成框架
  • 支持复杂特征交互与特定分布,适配多种实验需求
  • 适合评估推荐系统、检测算法偏见等研究场景

合成数据集对机器学习模型的评估至关重要。在评估真实推荐系统时,常需处理高维分类(且稀疏)数据集,但现有方法难以生成具备此类特性的数据。为此,我们提出一种新型框架,可生成多样且统计一致的合成数据集。该框架支持可控属性,允许迭代调整以满足特定实验需求,如引入复杂特征交互、特征基数或特定分布。通过基准测试概率计数算法、检测算法偏见和模拟AutoML搜索等用例,验证了其有效性。相比仅聚焦特定结构或依赖真实数据私有合成的方法,本框架提供模块化手段,快速生成完全合成数据,灵活适配多样实验。结果表明,该框架能有效隔离模型在特殊情境下的行为,展现出显著提升推荐系统评估与开发潜力。开源的Python工具包已发布,便于研究者低成本使用。

原文摘要 · Abstract (English)

Synthetic datasets are important for evaluating and testing machine learning models. When evaluating real-life recommender systems, high-dimensional categorical (and sparse) datasets are often considered. Unfortunately, there are not many solutions that would allow generation of artificial datasets with such characteristics. For that purpose, we developed a novel framework for generating synthetic datasets that are diverse and statistically coherent. Our framework allows for creation of datasets with controlled attributes, enabling iterative modifications to fit specific experimental needs, such as introducing complex feature interactions, feature cardinality, or specific distributions. We demonstrate the framework's utility through use cases such as benchmarking probabilistic counting algorithms, detecting algorithmic bias, and simulating AutoML searches. Unlike existing methods that either focus narrowly on specific dataset structures, or prioritize (private) data synthesis through real data, our approach provides a modular means to quickly generating completely synthetic datasets we can tailor to diverse experimental requirements. Our results show that the framework effectively isolates model behavior in unique situations and highlights its potential for significant advancements in the evaluation and development of recommender systems. The readily-available framework is available as a free open Python package to facilitate research with minimal friction.

合成数据推荐系统评估框架数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。