构建多领域多数据集基准,评估大模型推荐能力
Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- 构建跨领域、跨数据集的评估框架,支持零资源场景
- 19个大模型在15个数据集上测试,发现领域微调最优
- 多领域训练提升模型适应性,适合推荐系统研究者
现有基础模型在推荐任务中的能力评估亟需系统化。本文提出RecBench-MD,一个涵盖10个不同领域(包括电商、娱乐和社交媒体)的综合性评测基准,用于从零资源、多数据集、多域视角评估基础模型的推荐能力。我们在15个数据集上对19个基础模型进行了广泛评估,结果表明:领域内微调能取得最佳性能;跨数据集迁移学习为新推荐场景提供有效支持;多领域训练显著增强模型适应性。所有代码与数据均已公开,以推动后续研究。
原文摘要 · Abstract (English)
Comprehensive evaluation of the recommendation capabilities of existing foundation models across diverse datasets and domains is essential for advancing the development of recommendation foundation models. In this study, we introduce RecBench-MD, a novel and comprehensive benchmark designed to assess the recommendation abilities of foundation models from a zero-resource, multi-dataset, and multi-domain perspective. Through extensive evaluations of 19 foundation models across 15 datasets spanning 10 diverse domains -- including e-commerce, entertainment, and social media -- we identify key characteristics of these models in recommendation tasks. Our findings suggest that in-domain fine-tuning achieves optimal performance, while cross-dataset transfer learning provides effective practical support for new recommendation scenarios. Additionally, we observe that multi-domain training significantly enhances the adaptability of foundation models. All code and data have been publicly released to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。