arXiv:2409.15857cs.IR2024-09中稿 · Expert Systems wit…被引 6

首个大规模多模态推荐基准,系统评估特征提取效果

Large-scale Benchmarks for Multimodal Recommendation with Ducho

  • 构建统一实验环境,对比三种主流多模态提取框架
  • 在多领域、多模态数据上验证结果,发现提取器影响显著
  • 为下一代多模态推荐算法提供可复现的训练与调参参考

多模态推荐通常包括:(i) 提取多模态特征,(ii) 优化其高层表示以适配推荐任务,(iii) 可选地融合所有多模态特征,(iv) 预测用户-物品评分。尽管对(二)-(四)已有大量优化研究,但对(一)的系统性探索仍不足。现有文献虽有丰富多模态数据集和大型模型,却常盲目采用有限标准化方案。近期研究开始实证分析多模态对推荐的贡献,本文延续此方向,首次提出大规模多模态推荐基准,聚焦多模态提取器。我们基于Ducho、MMRec/Elliot三个主流框架,构建统一、可复现的实验环境,支持使用新型多模态特征提取器进行广泛基准测试。结果在不同提取器、超参数、领域和模态下均被验证,揭示了训练与调参新世代多模态推荐算法的关键洞见。

原文摘要 · Abstract (English)

The common multimodal recommendation pipeline involves (i) extracting multimodal features, (ii) refining their high-level representations to suit the recommendation task, (iii) optionally fusing all multimodal features, and (iv) predicting the user-item score. Although great effort has been put into designing optimal solutions for (ii-iv), to the best of our knowledge, very little attention has been devoted to exploring procedures for (i) in a rigorous way. In this respect, the existing literature outlines the large availability of multimodal datasets and the ever-growing number of large models accounting for multimodal-aware tasks, but (at the same time) an unjustified adoption of limited standardized solutions. As very recent works from the literature have begun to conduct empirical studies to assess the contribution of multimodality in recommendation, we decide to follow and complement this same research direction. To this end, this paper settles as the first attempt to offer a large-scale benchmarking for multimodal recommender systems, with a specific focus on multimodal extractors. Specifically, we take advantage of three popular and recent frameworks for multimodal feature extraction and reproducibility in recommendation, Ducho, and MMRec/Elliot, respectively, to offer a unified and ready-to-use experimental environment able to run extensive benchmarking analyses leveraging novel multimodal feature extractors. Results, largely validated under different extractors, hyper-parameters of the extractors, domains, and modalities, provide important insights on how to train and tune the next generation of multimodal recommendation algorithms.

多模态推荐特征提取基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。