arXiv:2504.10150cs.IRcs.MM2025-04被引 2

将用户多模态历史压缩为单个标记,提升大模型推荐效率与准确率。

HistLLM: A Unified Framework for LLM-Based Multimodal Recommendation with User History Encoding and Compression

  • 用编码模块将图文历史压缩成单一标记
  • 在多个数据集上显著提升推荐效果
  • 适合需要高效处理用户历史的推荐系统开发者

尽管大语言模型在文本推荐中表现优异,但在多模态推荐任务中应用仍较有限。虽然可通过投影函数将视觉特征映射至语义空间,但传统方法需拼接长篇图文交互历史作为提示,导致训练与推理效率低,且难以精准捕捉用户偏好,影响推荐性能。为此,我们提出HistLLM框架,通过用户历史编码模块(UHEM)融合文本与视觉特征,将多模态用户历史压缩为单个标记表示,有效支持大模型对用户偏好的建模。大量实验验证了该机制在效果与效率上的优越性。

原文摘要 · Abstract (English)

While large language models (LLMs) have proven effective in leveraging textual data for recommendations, their application to multimodal recommendation tasks remains relatively underexplored. Although LLMs can process multimodal information through projection functions that map visual features into their semantic space, recommendation tasks often require representing users' history interactions through lengthy prompts combining text and visual elements, which not only hampers training and inference efficiency but also makes it difficult for the model to accurately capture user preferences from complex and extended prompts, leading to reduced recommendation performance. To address this challenge, we introduce HistLLM, an innovative multimodal recommendation framework that integrates textual and visual features through a User History Encoding Module (UHEM), compressing multimodal user history interactions into a single token representation, effectively facilitating LLMs in processing user preferences. Extensive experiments demonstrate the effectiveness and efficiency of our proposed mechanism.

多模态推荐大模型用户历史压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。