arXiv:2508.05377cs.IRcs.MM2025-08被引 4

多模态推荐并非总有效,关键看场景和阶段。

Does Multimodality Improve Recommender Systems as Expected? A Critical Analysis and Future Directions

  • 构建四维度评估框架,系统测试多模态推荐效果
  • 稀疏数据与召回阶段最受益,图文/视觉模态各有所长
  • 集成学习优于融合学习,大模型未必更优

多模态推荐系统因整合多元数据而日益流行,但其实际增益尚不明确。本文提出一个结构化评估框架,从对比效率、推荐任务、推荐阶段和多模态融合四个维度系统评估。在多个平台对可复现的多模态模型与强基线进行基准测试。结果表明:多模态数据在交互稀疏场景和召回阶段尤为有效;文本特征在电商中更优,视觉特征在短视频推荐中表现更好;集成学习优于融合学习,且模型越大并不一定越好。通过案例研究和跨领域回顾,本文为构建高效多模态推荐系统提供实践指导,强调需谨慎选择模态、融合策略与模型设计。

原文摘要 · Abstract (English)

Multimodal recommendation systems are increasingly popular for their potential to improve performance by integrating diverse data types. However, the actual benefits of this integration remain unclear, raising questions about when and how it truly enhances recommendations. In this paper, we propose a structured evaluation framework to systematically assess multimodal recommendations across four dimensions: Comparative Efficiency, Recommendation Tasks, Recommendation Stages, and Multimodal Data Integration. We benchmark a set of reproducible multimodal models against strong traditional baselines and evaluate their performance on different platforms. Our findings show that multimodal data is particularly beneficial in sparse interaction scenarios and during the recall stage of recommendation pipelines. We also observe that the importance of each modality is task-specific, where text features are more useful in e-commerce and visual features are more effective in short-video recommendations. Additionally, we explore different integration strategies and model sizes, finding that Ensemble-Based Learning outperforms Fusion-Based Learning, and that larger models do not necessarily deliver better results. To deepen our understanding, we include case studies and review findings from other recommendation domains. Our work provides practical insights for building efficient and effective multimodal recommendation systems, emphasizing the need for thoughtful modality selection, integration strategies, and model design.

推荐系统多模态评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。