arXiv:2601.11182cs.IR2026-01被引 2

用稀疏自编码器让推荐系统可解释、可控制。

From Knots to Knobs: Towards Steerable Collaborative Filtering Using Sparse Autoencoders

  • 在协同过滤模型中插入稀疏自编码器,提取可解释特征。
  • 发现隐层神经元基本单一语义,能对应具体推荐概念。
  • 可通过激活特定神经元精准调整推荐结果,适合可解释推荐场景。

稀疏自编码器(SAEs)近期成为剖析大语言模型的重要工具,能够揭示多层次、高可解释性的特征,并通过选择性激活潜在空间中的特定神经元实现生成过程的定向调控。本文首次将该方法应用于协同过滤领域,旨在从纯交互信号学习的表示中提取类似可解释的特征。我们聚焦于一类广泛应用的协同自编码器(CFAE),在其编码器与解码器之间插入一个SAE。实验表明,该表示具有高度单义性,并提出了语义概念与单个神经元之间的合适映射函数。此外,我们评估了一种简单而有效的方法,利用该表示实现对推荐结果的定向引导。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have recently emerged as pivotal tools for introspection into large language models. SAEs can uncover high-quality, interpretable features at different levels of granularity and enable targeted steering of the generation process by selectively activating specific neurons in their latent activations. Our paper is the first to apply this approach to collaborative filtering, aiming to extract similarly interpretable features from representations learned purely from interaction signals. In particular, we focus on a widely adopted class of collaborative autoencoders (CFAEs) and augment them by inserting an SAE between their encoder and decoder networks. We demonstrate that such representation is largely monosemantic and propose suitable mapping functions between semantic concepts and individual neurons. We also evaluate a simple yet effective method that utilizes this representation to steer the recommendations in a desired direction.

协同过滤可解释性稀疏自编码器推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。