arXiv:2502.15542cs.IRcs.AI2025-02被引 1

用高效微调解决视觉语言模型与推荐系统间的领域差异问题

Bridging Domain Gaps between Pretrained Multimodal Models and Recommendations

  • 设计双阶段参数高效训练,引导知识迁移以缩小领域差距
  • 在不额外预训练的情况下,性能超越基线模型
  • 支持多种高效微调方法,兼顾效果与计算效率

随着在线多模态内容的爆炸式增长,预训练视觉-语言模型在多模态推荐中展现出巨大潜力。然而,尽管冻结使用时表现良好,但由于预训练与个性化推荐之间存在显著领域差距(如特征分布差异和任务目标错配),采用联合训练反而导致性能低于基线。现有方法或依赖简单特征提取,或需昂贵的全模型微调,难以兼顾效果与效率。为此,我们提出参数高效多模态推荐框架PTMRec,通过知识引导的双阶段参数高效训练策略,有效弥合预训练模型与推荐系统之间的领域差距。该框架无需额外预训练,且可灵活适配多种参数高效微调方法。

原文摘要 · Abstract (English)

With the explosive growth of multimodal content online, pre-trained visual-language models have shown great potential for multimodal recommendation. However, while these models achieve decent performance when applied in a frozen manner, surprisingly, due to significant domain gaps (e.g., feature distribution discrepancy and task objective misalignment) between pre-training and personalized recommendation, adopting a joint training approach instead leads to performance worse than baseline. Existing approaches either rely on simple feature extraction or require computationally expensive full model fine-tuning, struggling to balance effectiveness and efficiency. To tackle these challenges, we propose \textbf{P}arameter-efficient \textbf{T}uning for \textbf{M}ultimodal \textbf{Rec}ommendation (\textbf{PTMRec}), a novel framework that bridges the domain gap between pre-trained models and recommendation systems through a knowledge-guided dual-stage parameter-efficient training strategy. This framework not only eliminates the need for costly additional pre-training but also flexibly accommodates various parameter-efficient tuning methods.

多模态推荐参数高效视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。