arXiv:2507.12202cs.IRcs.AI2025-07被引 3

用稀疏自编码器让推荐模型更透明,还能自由调控推荐结果。

Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control

  • 用无监督稀疏自编码器提取可解释的推荐特征。
  • 学习到的方向比原隐藏层更清晰、单一语义。
  • 开发者可灵活控制推荐行为,适应不同场景。

当前主流的序列推荐模型多基于Transformer架构,其黑箱特性阻碍了对模型内部机制的理解,而理解与控制模型行为在真实应用场景中至关重要。近期研究表明,稀疏自编码器(SAE)是一种有前景的无监督方法,可用于从神经网络中提取可解释特征。本文将SAE扩展至序列推荐系统,提出一个用于解释和控制模型表征的框架。实验表明,该方法可成功应用于训练于序列推荐任务的Transformer模型:在无监督条件下学习到的方向比原始隐藏状态维度更具可解释性且语义单一。此外,我们展示了一种简单有效的方式,实现对模型行为的灵活控制,使推荐系统的开发者和用户能够根据具体需求调整推荐结果。

原文摘要 · Abstract (English)

Many current state-of-the-art models for sequential recommendations are based on transformer architectures. Interpretation and explanation of such black box models is an important research question, as a better understanding of their internals can help understand, influence, and control their behavior, which is very important in a variety of real-world applications. Recently, sparse autoencoders (SAE) have been shown to be a promising unsupervised approach to extract interpretable features from neural networks. In this work, we extend SAE to sequential recommender systems and propose a framework for interpreting and controlling model representations. We show that this approach can be successfully applied to the transformer trained on a sequential recommendation task: directions learned in such an unsupervised regime turn out to be more interpretable and monosemantic than the original hidden state dimensions. Further, we demonstrate a straightforward way to effectively and flexibly control the model's behavior, giving developers and users of recommendation systems the ability to adjust their recommendations to various custom scenarios and contexts.

推荐系统可解释性稀疏编码控制模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。