用稀疏潜在特征解释并控制稠密检索向量的输出
Interpret and Control Dense Retrieval with Sparse Latent Features
- 通过稀疏自编码器提取可解释的潜在特征
- 重构后的向量与原始稠密向量检索精度几乎一致
- 可操控潜在特征实现特定视角的检索结果
稠密嵌入虽具强检索性能,但缺乏可解释性与可控性。本文提出一种新方法,利用稀疏自编码器(SAE)通过学习到的稀疏潜在特征来解释和控制稠密嵌入。关键贡献在于设计了一种面向检索的对比损失,确保稀疏潜在特征在检索任务中仍保持有效性,从而具备可解释意义。实验表明,所学稀疏潜在特征及其重构嵌入的检索准确率几乎与原始稠密向量持平,验证了其忠实性。进一步分析显示,稀疏潜在空间揭示了稠密嵌入背后的有趣特征,且可通过操纵潜在特征控制检索行为,例如优先返回特定视角的文档。
原文摘要 · Abstract (English)
Dense embeddings deliver strong retrieval performance but often lack interpretability and controllability. This paper introduces a novel approach using sparse autoencoders (SAE) to interpret and control dense embeddings via the learned latent sparse features. Our key contribution is the development of a retrieval-oriented contrastive loss, which ensures the sparse latent features remain effective for retrieval tasks and thus meaningful to interpret. Experimental results demonstrate that both the learned latent sparse features and their reconstructed embeddings retain nearly the same retrieval accuracy as the original dense vectors, affirming their faithfulness. Our further examination of the sparse latent space reveals interesting features underlying the dense embeddings and we can control the retrieval behaviors via manipulating the latent sparse features, for example, prioritizing documents from specific perspectives in the retrieval results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。