arXiv:2504.15085cs.CV2025-04中稿 · CogSCI 2025被引 4

用视觉文本融合模拟人类认知,提升跨域序列推荐效果

Cognitive-Inspired Hierarchical Attention Fusion With Visual and Textual for Cross-Domain Sequential Recommendation

  • 分层注意力融合图像与文本特征,模拟人类信息整合过程
  • 在四个电商数据集上优于现有方法,显著提升跨域兴趣捕捉能力
  • 适合研究多模态推荐、认知计算或跨域行为建模的读者

跨域序列推荐(CDSR)通过利用多领域历史交互预测用户行为,重点在于通过域内与域间项目关系建模跨域偏好。受人类认知过程启发,本文提出视觉与文本表示的分层注意力融合(HAF-VT)方法,整合多模态数据以增强认知建模。采用冻结的CLIP模型生成图像与文本嵌入,丰富项目表征。分层注意力机制联合学习单域与跨域偏好,模拟人类信息整合过程。在四个电商平台数据集上的实验表明,HAF-VT在捕捉跨域用户兴趣方面优于现有方法,将认知原理与计算模型结合,凸显多模态数据在序列决策中的作用。

原文摘要 · Abstract (English)

Cross-Domain Sequential Recommendation (CDSR) predicts user behavior by leveraging historical interactions across multiple domains, focusing on modeling cross-domain preferences through intra- and inter-sequence item relationships. Inspired by human cognitive processes, we propose Hierarchical Attention Fusion of Visual and Textual Representations (HAF-VT), a novel approach integrating visual and textual data to enhance cognitive modeling. Using the frozen CLIP model, we generate image and text embeddings, enriching item representations with multimodal data. A hierarchical attention mechanism jointly learns single-domain and cross-domain preferences, mimicking human information integration. Evaluated on four e-commerce datasets, HAF-VT outperforms existing methods in capturing cross-domain user interests, bridging cognitive principles with computational models and highlighting the role of multimodal data in sequential decision-making.

跨域推荐多模态融合认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。