研究发现,文本评论对推荐系统提升有限,协同信号仍占主导。
How Much Do Reviews Really Contribute? A Study on Text-Enriched Matrix Factorization for Recommendations
- 设计可学习门控机制,动态融合协同与文本信号。
- 在多个数据集上,文本信息的增益远低于协同信号。
- 适合关注评论如何有效融入推荐系统的研究人员。
将文本评论融入推荐系统已成为利用语义信息增强协同信号的主流策略。然而,在强协同基线条件下,评论衍生表示的实际贡献仍不明确。本文通过引入并比较三种基于统一协同主干的增强策略,系统研究了文本信息对矩阵分解的影响。首先提出可学习门控机制,在训练中自适应平衡协同与文本信号;该机制应用于两种不同评论表示:(i) 从用户和物品历史中提取的聚合主题特征,(ii) 由评论生成的完整文本嵌入。此外,还探索了跨注意力机制,用于识别并强化文本表示中最关键的维度。共评估六种变体:纯协同模型,以及分别使用主题和文本通过门控融合、或通过跨注意力增强的模型。在多个基于评论的数据集上实验表明,尽管自适应融合提升了表示灵活性,但相比协同主干,文本信号的边际贡献仍有限。结果表明,在典型评分预测任务中,协同信息仍占据主导地位,提示需重新思考语义评论信号的有效整合方式。
原文摘要 · Abstract (English)
Incorporating textual reviews into a Recommender System has become a prominent strategy for enriching collaborative signals with semantic information. However, the actual contribution of review-derived representations remains an open question, particularly when strong collaborative baselines are employed. In this work, we systematically investigate the impact of textual information on Matrix Factorization by introducing and comparing three enrichment strategies over a common collaborative backbone. First, we propose a learnable gating mechanism that adaptively balances collaborative and textual signals during training. This mechanism is applied to two distinct review representations: (i) aggregated topic profiles extracted from user and item histories, and (ii) full text embedding representations derived from reviews. Additionally, we explore a cross-attention mechanism that identifies and emphasizes the most informative dimensions of the textual representation before fusion with collaborative factors. We evaluate six variants: pure, enriched with topic profiles and text via gating; enriched with topics and text via gating; and enhanced with cross-attention over textual features. Experiments across multiple review-based datasets reveal that although adaptive fusion mechanisms improve representation flexibility, the marginal contribution of textual signals remains limited compared to the collaborative backbone. These findings suggest that, under typical rating-prediction settings, collaborative information continues to dominate performance, raising important considerations for the effective integration of semantic review signals into recommendation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。