用人格向量构建低维空间,提升模型行为探测的泛化能力
Do Linear Probes Generalize Better in Persona Coordinates?

- 基于对比人格提示构造人格轴,提取有害行为的主成分方向
- 在10个数据集上,基于人格主成分的探测器泛化性能显著优于原始激活值探测
- 统一有害与无害行为的综合轴线,进一步增强跨任务和跨数据集的适应性
语言模型交互中检测有害行为日益重要,但仅靠文本监测难以应对模型策略性欺骗与藏拙行为。为此,需采用白盒监控手段如线性探测器,直接读取模型内部状态。然而现有探测器在分布外场景下表现不佳。本文受助理轴与人格选择模型启发,利用对比人格提示构建欺骗与谄媚的人格轴。通过无监督PCA提取人格特定向量的第一主成分,能清晰区分有害与无害人格。在10个评估数据集上,基于人格主成分投影的探测器展现出非平凡的迁移能力,其泛化性能优于基于原始激活值训练的探测器。同时发现,融合多种有害与无害行为的统一轴线可进一步提升跨行为、跨数据集的泛化效果。总体而言,人格向量为构建更具迁移性的行为探测器提供了有效归纳偏置。
原文摘要 · Abstract (English)
It is becoming increasingly necessary to have monitors check for harmful behaviors during language model interactions, but text-only monitoring has not been sufficient. This is because models sometimes exhibit strategic deception and sandbagging, changing their behavior during evaluation. This motivates the use of white-box monitors like linear probes, which can read the model internals directly. Currently, such probes can fail under distribution shift, limiting their usefulness in real settings. We study whether there exists a low-dimensional subspace of the model internals that captures harmful behaviors more robustly, while leaving out spuriously correlative features. Inspired by the Assistant Axis and Persona Selection Model, we construct persona axes for deception and sycophancy using contrastive persona prompts. The first principal components, obtained by unsupervised PCA of the persona-specific vectors, cleanly separate harmful and harmless personas. Across 10 evaluation datasets, we show that persona-derived directions transfer non-trivially and probes trained on persona-PC projections generalize better than probes trained on raw activations. We also find that a unified axis consisting of multiple harmful and harmless behaviors improves generalization across behaviors and datasets. Overall, persona vectors provide a useful inductive bias for building more transferable behavior probes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。