arXiv:2504.10307cs.IR2025-04被引 5

多模态大模型高效适配序列推荐,训练快省显存

CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation

  • 用解耦侧边适配器实现多模型高效协同
  • 训练时间降30%,参数少57%,性能提6.7%~8.1%
  • 适合需融合多个模态大模型的推荐系统研究

本文探讨了一个较少研究但实际重要的问题:如何高效且有效地将多个(>2)多模态基础模型(MFMs)适配到序列推荐任务中。为此,我们提出一种即插即用的跨模态侧边适配网络(CROSSAN),采用完全解耦的侧边适配器范式,实现高效可扩展的适配。相比现有最先进方法,CROSSAN将训练时间减少超过30%,GPU内存消耗降低20%,可训练参数减少超过57%,同时支持跨模态的有效学习。为进一步提升多模态融合效果,我们引入了模态专家混合融合机制(MOMEF)。在公开基准上的大量实验表明,当适配四个基础模型且使用原始模态时,CROSSAN性能提升6.7%至8.1%。随着更多MFMs的引入,整体性能持续提升。代码与数据集将公开,以促进后续研究。

原文摘要 · Abstract (English)

In this paper, we explore a less-studied yet practically important problem: how to efficiently and effectively adapt multiple ($>$2) multimodal foundation models (MFMs) for the sequential recommendation task. To this end, we propose a plug-and-play Cross-modal Side Adapter Network (CROSSAN), which leverages a fully decoupled side adapter-based paradigm to achieve efficient and scalable adaptation. Compared to the state-of-the-art efficient approaches, CROSSAN reduces training time by over 30%, GPU memory consumption by 20%, and trainable parameters by over 57%, while enabling effective cross-modal learning across diverse modalities. To further enhance multimodal fusion, we introduce the Mixture of Modality Expert Fusion (MOMEF) mechanism. Extensive experiments on public benchmarks demonstrate that CROSSAN consistently outperforms existing methods, achieving 6.7%--8.1% performance improvements when adapting four foundation models with raw modalities. Moreover, the overall performance continues to improve as more MFMs are incorporated. We will release our code and datasets to faciliate future research.

序列推荐多模态高效适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。