arXiv:2502.05407cs.LGcs.AI2025-02

用语言模型反馈高效学习稀疏特征,理论与实验结合。

Dictionary Learning: The Complexity of Learning Sparse Superposed Features with Feedback

  • 通过三元组比较反馈,从稀疏特征中恢复潜在表示
  • 在稀疏场景下,反馈仅需分布信息也能实现强上界
  • 适用于大模型特征提取与稀疏自编码器研究

深度网络的成功关键在于其能捕捉表示空间中的潜在特征。本文研究能否通过智能体(如大语言模型)的相对三元组比较反馈,高效还原模型所学特征。这些特征可能对应语言模型中的字典或马氏距离协方差矩阵。我们分析了在稀疏设置下学习特征矩阵时的反馈复杂度。当智能体可构造激活时,建立了紧致的复杂度边界;当反馈仅限于分布信息时,在稀疏场景下仍获得强上界。通过两个应用验证:从递归特征机中恢复特征,以及从训练于大语言模型的稀疏自编码器中提取字典。

原文摘要 · Abstract (English)

The success of deep networks is crucially attributed to their ability to capture latent features within a representation space. In this work, we investigate whether the underlying learned features of a model can be efficiently retrieved through feedback from an agent, such as a large language model (LLM), in the form of relative \tt{triplet comparisons}. These features may represent various constructs, including dictionaries in LLMs or a covariance matrix of Mahalanobis distances. We analyze the feedback complexity associated with learning a feature matrix in sparse settings. Our results establish tight bounds when the agent is permitted to construct activations and demonstrate strong upper bounds in sparse scenarios when the agent's feedback is limited to distributional information. We validate our theoretical findings through experiments on two distinct applications: feature recovery from Recursive Feature Machines and dictionary extraction from sparse autoencoders trained on Large Language Models.

特征学习稀疏表示大模型反馈机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。