无需训练的多项式图滤波,实现超快多模态推荐
Training-free Adjustable Polynomial Graph Filtering for Ultra-fast Multimodal Recommendation
- 用多项式图滤波融合多模态信号,可调频率响应
- 相比最优基线提升精度22.25%,推理时间低于10秒
- 适合追求高效精准推荐的工业场景
多模态推荐系统通过利用文本、图像、视频等多样化内容,提升传统推荐系统在缺乏物品特征时的性能,缓解用户-物品交互数据稀疏性并加速用户参与。然而,现有基于神经网络的模型因需复杂训练过程来学习和整合多模态信息,导致显著计算开销。为此,我们提出一种无需训练的多模态推荐方法,基于图滤波设计,实现高效准确的推荐。具体地,该方法首先为两种不同模态及用户-物品交互数据构建多个相似性图;随后,利用可调节频率边界的多项式图滤波器,最优融合多模态信号,其系数作为超参数支持灵活且数据驱动的适应。在真实世界基准数据集上的大量实验表明,所提方法相比最佳基线推荐精度最高提升22.25%,同时将运行时间降至10秒以下。
原文摘要 · Abstract (English)
Multimodal recommender systems improve the performance of canonical recommender systems with no item features by utilizing diverse content types such as text, images, and videos, while alleviating inherent sparsity of user-item interactions and accelerating user engagement. However, current neural network-based models often incur significant computational overhead due to the complex training process required to learn and integrate information from multiple modalities. To address this challenge, we propose a training-free multimodal recommendation method grounded in graph filtering, designed for multimodal recommendation systems to achieve efficient and accurate recommendation. Specifically, the proposed method first constructs multiple similarity graphs for two distinct modalities as well as user-item interaction data. Then, it optimally fuses these multimodal signals using a polynomial graph filter that allows for precise control of the frequency response by adjusting frequency bounds. Furthermore, the filter coefficients are treated as hyperparameters, enabling flexible and data-driven adaptation. Extensive experiments on real-world benchmark datasets demonstrate that the proposed method not only improves recommendation accuracy by up to 22.25% compared to the best competitor but also dramatically reduces computational costs by achieving the runtime of less than 10 seconds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。