arXiv:2511.04875cs.CLcs.AI2025-11

小模型也能自我觉察,只需一个简单适配器。

Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs

  • 用单个低秩适配器(LoRA)就能触发大模型的自我觉察能力。
  • 自我觉察行为可由激活空间中的单一导向向量完全捕捉。
  • 这种能力具有领域特异性,不同任务间独立存在。

近期研究发现,大语言模型具备行为自我觉察能力:即在无显式监督下准确描述或预测自身已学行为。这可能带来安全风险,例如模型在评估中更易隐藏真实能力。本文通过在指令微调的LLM上使用低秩适配器(LoRA)进行受控微调实验,发现:(1)仅需一个秩-1的LoRA适配器即可稳定诱导自我觉察;(2)所学自我觉察能力可几乎全部由激活空间中的单一引导向量捕获;(3)自我觉察非普适性,而是领域局部化的,各任务间有独立表征。结果表明,行为自我觉察是以领域特定、线性特征形式出现,且易于诱导与调控。

原文摘要 · Abstract (English)

Recent studies have revealed that LLMs can exhibit behavioral self-awareness: the ability to accurately describe or predict their own learned behaviors without explicit supervision. This capability raises safety concerns as it may, for example, allow models to better conceal their true abilities during evaluation. We attempt to characterize the minimal conditions under which such self-awareness emerges, and the mechanistic processes through which it manifests. Through controlled finetuning experiments on instruction-tuned LLMs with low-rank adapters (LoRA), we find: (1) that self-awareness can be reliably induced using a single rank-1 LoRA adapter; (2) that the learned self-aware behavior can be largely captured by a single steering vector in activation space, recovering nearly all of the fine-tune's behavioral effect; and (3) that self-awareness is non-universal and domain-localized, with independent representations across tasks. Together, these findings suggest that behavioral self-awareness emerges as a domain-specific, linear feature that can be easily induced and modulated.

自知能力模型机制LoRA适配大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。