通过双图学习与跨模态注意力,提升多模态推荐的融合精度与用户建模对称性。
Cross-Modal Attention Network with Dual Graph Learning in Multimodal Recommendation
- 设计递归跨模态注意力机制,迭代挖掘多模态间高阶依赖关系。
- 构建对称双图框架,融合行为与语义信号,实现用户与物品同等建模。
- 在四个真实数据集上平均提升5%关键指标,兼具高效与可扩展性。
多模态推荐系统利用用户-物品交互和多模态信息捕捉用户偏好,实现更精准个性化的推荐。尽管取得显著进展,现有方法仍存在两大局限:其一,浅层模态融合常依赖简单拼接,难以挖掘丰富的模态内与模态间协同关系;其二,特征处理不对称——用户仅由交互ID表征,而物品则受益于丰富的多模态内容,阻碍了共享语义空间的学习。为此,我们提出交叉模态递归注意力网络与双图嵌入(CRANE)。为解决融合浅层问题,设计核心递归跨模态注意力(RCA)机制,基于联合潜在空间中的交叉相关性迭代优化模态特征,有效捕捉高阶模态内与模态间依赖。针对对称多模态学习,显式通过聚合用户交互物品特征构建用户多模态画像。此外,CRANE集成对称双图框架——异质用户-物品交互图与同质物品-物品语义图——通过自监督对比学习目标统一行为与语义信号。尽管具备复杂建模能力,CRANE保持高计算效率。理论与实证分析证实其可扩展性与高实用性,在小数据集上收敛更快,在大规模数据集上性能上限更高。在四个公开真实世界数据集上的综合实验验证,其关键指标平均优于当前最优基线5%。
原文摘要 · Abstract (English)
Multimedia recommendation systems leverage user-item interactions and multimodal information to capture user preferences, enabling more accurate and personalized recommendations. Despite notable advancements, existing approaches still face two critical limitations: first, shallow modality fusion often relies on simple concatenation, failing to exploit rich synergic intra- and inter-modal relationships; second, asymmetric feature treatment-where users are only characterized by interaction IDs while items benefit from rich multimodal content-hinders the learning of a shared semantic space. To address these issues, we propose a Cross-modal Recursive Attention Network with dual graph Embedding (CRANE). To tackle shallow fusion, we design a core Recursive Cross-Modal Attention (RCA) mechanism that iteratively refines modality features based on cross-correlations in a joint latent space, effectively capturing high-order intra- and inter-modal dependencies. For symmetric multimodal learning, we explicitly construct users' multimodal profiles by aggregating features of their interacted items. Furthermore, CRANE integrates a symmetric dual-graph framework-comprising a heterogeneous user-item interaction graph and a homogeneous item-item semantic graph-unified by a self-supervised contrastive learning objective to fuse behavioral and semantic signals. Despite these complex modeling capabilities, CRANE maintains high computational efficiency. Theoretical and empirical analyses confirm its scalability and high practical efficiency, achieving faster convergence on small datasets and superior performance ceilings on large-scale ones. Comprehensive experiments on four public real-world datasets validate an average 5% improvement in key metrics over state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。