不同任务间正确答案的激活模式相互正交,无法通用。
The Geometries of Truth Are Orthogonal Across Tasks
- 用线性分类器区分正确与错误回答的激活模式
- 跨任务训练的分类器相似度极低,支持集几乎不重叠
- 适合研究模型可靠性与任务间泛化局限的学者
大语言模型在多任务上表现出优异的泛化能力,但其实际可靠性仍受质疑。近期研究尝试通过分析推理时的激活值来判断答案是否正确,认为可通过学习‘真理几何’——即正确回答与错误回答的激活模式可被线性分类器区分。本文揭示该方法的关键局限:这些‘真理几何’具有强任务依赖性,无法跨任务迁移。我们发现,不同任务间训练的线性分类器之间相似度极低,且在稀疏正则化下支持集几乎完全分离。更复杂的混合探针或任务联合方法也无法克服此问题,原因在于各类任务的激活向量在空间中形成明显分离的聚类。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated impressive generalization capabilities across various tasks, but their claim to practical relevance is still mired by concerns on their reliability. Recent works have proposed examining the activations produced by an LLM at inference time to assess whether its answer to a question is correct. Some works claim that a "geometry of truth" can be learned from examples, in the sense that the activations that generate correct answers can be distinguished from those leading to mistakes with a linear classifier. In this work, we underline a limitation of these approaches: we observe that these "geometries of truth" are intrinsically task-dependent and fail to transfer across tasks. More precisely, we show that linear classifiers trained across distinct tasks share little similarity and, when trained with sparsity-enforcing regularizers, have almost disjoint supports. We show that more sophisticated approaches (e.g., using mixtures of probes and tasks) fail to overcome this limitation, likely because activation vectors commonly used to classify answers form clearly separated clusters when examined across tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。