arXiv:2506.00041cs.IRcs.LG2025-06EMNLP被引 10

用稀疏自编码器拆解稠密检索的黑箱,让模型可解释且更高效

Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrieval

  • 用稀疏自编码器分解稠密向量,提取可读的语义概念
  • 新检索框架在多种不匹配场景下仍保持高效率与高性能
  • 适合想理解检索模型、提升可解释性的研究人员

尽管密集段落检索(DPR)模型表现优异,但缺乏可解释性。本文提出一种新型可解释性框架,利用稀疏自编码器(SAEs)将原先不可解释的DPR稠密嵌入分解为独立、可解释的潜在语义概念,并为每个概念生成自然语言描述,使人类能理解嵌入和查询-文档相似度。我们进一步提出概念级稀疏检索(CL-SR),直接以提取出的潜在概念作为索引单元。该框架结合了稠密表示的语义表达力与稀疏表示的透明性与高效性,在词汇与语义不匹配场景下仍具备高索引空间与计算效率,同时保持稳健性能。

原文摘要 · Abstract (English)

Despite their strong performance, Dense Passage Retrieval (DPR) models suffer from a lack of interpretability. In this work, we propose a novel interpretability framework that leverages Sparse Autoencoders (SAEs) to decompose previously uninterpretable dense embeddings from DPR models into distinct, interpretable latent concepts. We generate natural language descriptions for each latent concept, enabling human interpretations of both the dense embeddings and the query-document similarity scores of DPR models. We further introduce Concept-Level Sparse Retrieval (CL-SR), a retrieval framework that directly utilizes the extracted latent concepts as indexing units. CL-SR effectively combines the semantic expressiveness of dense embeddings with the transparency and efficiency of sparse representations. We show that CL-SR achieves high index-space and computational efficiency while maintaining robust performance across vocabulary and semantic mismatches.

可解释性稠密检索稀疏编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。