让推荐系统的隐向量变得可解释,发现性别相关特征可干预
Monosemanticity in Recommender Systems

- 用嵌套稀疏自编码器挖掘协同过滤嵌入的分层语义结构
- 在亚马逊时尚数据集上识别出具有明确语义的12个可解释特征
- 适合关注推荐系统可解释性与公平性的研究者阅读
矩阵分解等潜在因子模型广泛应用于推荐系统,但学习到的嵌入维度通常缺乏明确语义,限制了透明度、可解释性及可控干预。尽管稀疏自编码器(SAEs)可用于提取密集神经表示中的单义特征,但标准SAE存在特征分裂、吸收和组合等缩放病态,随字典规模增大而恶化可解释性。本文在亚马逊时尚数据集上训练大规模矩阵分解推荐系统,并应用马特约什卡稀疏自编码器(MSAE)分析学习到的嵌入。通过元数据对齐和大语言模型生成标签评估潜在特征的语义一致性与解耦性。最后,对分析中浮现的一组与性别相关的潜在神经元实施干预。结果表明,协同过滤嵌入中存在可恢复的分层结构,且马特约什卡训练提供了一种暴露交互驱动推荐模型中可解释潜在因子的合理机制。
原文摘要 · Abstract (English)
Latent factor models such as matrix factorization are widely used in recommender systems, yet the learned embedding dimensions typically lack explicit semantic interpretation. This opacity limits transparency, explainability, and principled intervention in recommendation behavior. While sparse autoencoders (SAEs) have recently been used to extract monosemantic features from dense neural representations, standard SAEs suffer from scaling pathologies including feature splitting, feature absorption, and feature composition, which degrade interpretability as dictionary size increases. In this work, we investigate whether hierarchical sparse representations can reveal interpretable structure in collaborative filtering embeddings. We train a large-scale matrix factorization recommender system on the Amazon Fashion dataset and apply a Matryoshka Sparse Autoencoder (MSAE) to the learned embeddings. We analyze the resulting latent features through metadata alignment and LLM-generated labeling to assess semantic coherence and disentanglement. Finally, we show an intervention on a subset of gender associated latent neurons that emerged from the analysis. Our findings suggest that collaborative filtering embeddings contain recoverable hierarchical structure, and that Matryoshka training provides a principled mechanism for exposing interpretable latent factors in interaction-driven recommendation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。