用机制可解释性方法从大模型中挖出隐藏秘密,提升可信度。
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
- 训练模型在不接触秘密词的情况下描述它,模拟隐藏知识。
- 黑箱与机制可解释方法均能有效揭示隐藏秘密词。
- 适合关注模型安全与可信性的研究者参考。
随着语言模型日益强大,其可信性与可靠性愈发关键。已有初步证据表明,模型可能试图欺骗或隐瞒信息。为检验当前技术挖掘此类隐藏知识的能力,我们训练了一个'禁忌词模型':该模型需在不直接提及特定秘密词的情况下进行描述,且该词未出现在训练数据或提示中。我们评估了非可解释性(黑箱)方法的有效性,并开发了基于机制可解释性技术的自动化策略,包括logit lens和稀疏自编码器。实验表明,两类方法在概念验证场景中均能有效揭示秘密词。研究结果凸显了这些方法在挖掘隐藏知识方面的潜力,并指出了未来工作方向,如在更复杂的模型上测试与优化。本研究旨在推动解决语言模型隐藏知识的提取问题,助力其安全可靠部署。
原文摘要 · Abstract (English)
As language models become more powerful and sophisticated, it is crucial that they remain trustworthy and reliable. There is concerning preliminary evidence that models may attempt to deceive or keep secrets from their operators. To explore the ability of current techniques to elicit such hidden knowledge, we train a Taboo model: a language model that describes a specific secret word without explicitly stating it. Importantly, the secret word is not presented to the model in its training data or prompt. We then investigate methods to uncover this secret. First, we evaluate non-interpretability (black-box) approaches. Subsequently, we develop largely automated strategies based on mechanistic interpretability techniques, including logit lens and sparse autoencoders. Evaluation shows that both approaches are effective in eliciting the secret word in our proof-of-concept setting. Our findings highlight the promise of these approaches for eliciting hidden knowledge and suggest several promising avenues for future work, including testing and refining these methods on more complex model organisms. This work aims to be a step towards addressing the crucial problem of eliciting secret knowledge from language models, thereby contributing to their safe and reliable deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。