arXiv:2411.06112cs.IR2024-11被引 4

用稀疏自编码器解析推荐模型内部表示,提升透明度与可解释性。

Understanding Internal Representations of Recommendation Models with Sparse Autoencoders

  • 通过稀疏自编码器从模型内部提取可解释的潜在特征。
  • 在4个数据集上验证了对三类推荐模型的有效性与通用性。
  • 结果获专家认可,适合需要模型透明化的研究与应用者。

推荐模型解释旨在揭示输入、内部表示与输出之间的关系,以增强推荐系统的透明性、可解释性和可信度。然而,深度学习模型固有的复杂性和不透明性给模型层面的解释带来挑战。此外,现有大多数推荐模型解释方法针对特定架构设计,限制了其在不同推荐模型间的泛化能力。本文提出RecSAE,一种通用的探测框架,利用稀疏自编码器解释推荐模型。该框架从推荐模型的内部表示中提取可解释的潜在变量,并将其与语义概念关联以实现解释。该方法在解释过程中不修改原始模型,同时支持对模型进行定向调优。在三类推荐模型(通用型、图结构型、序列型)和四个广泛使用的公开数据集上的实验表明,RecSAE框架具有有效性与良好的泛化能力。所识别的概念经由人类专家验证,与人类认知高度一致。总体而言,RecSAE为无需影响模型功能的各类推荐模型提供了一种新型的模型级解释方法,并具备潜在的模型定向调优价值。

原文摘要 · Abstract (English)

Recommendation model interpretation aims to reveal the relationships between inputs, model internal representations and outputs to enhance the transparency, interpretability, and trustworthiness of recommendation systems. However, the inherent complexity and opacity of deep learning models pose challenges for model-level interpretation. Moreover, most existing methods for interpreting recommendation models are tailored to specific architectures or model types, limiting their generalizability across different types of recommenders. In this paper, we propose RecSAE, a generalizable probing framework that interprets Recommendation models with Sparse AutoEncoders. The framework extracts interpretable latents from the internal representations of recommendation models, and links them to semantic concepts for interpretations. It does not alter original models during interpretations and also enables targeted tuning to models. Experiments on three types of recommendation models (general, graph-based, sequential) with four widely used public datasets demonstrate the effectiveness and generalization of RecSAE framework. The interpreted concepts are further validated by human experts, showing strong alignment with human perception. Overall, RecSAE serves as a novel step in both model-level interpretations to various types of recommendation models without affecting their functions and offering potential for targeted tuning of models.

推荐系统可解释性自编码器模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。