跨模型家族的行为轴具有一致性,可通过锚点投影实现高效迁移。
Cross-Family Universality of Behavioral Axes via Anchor-Projected Representations

- 用锚点投影将不同模型的隐藏表示映射到共享空间,统一行为方向
- 在LQMP模型簇中,行为轴对齐度高,下游检测准确率达0.83
- 仅需两个源模型和少量锚点即可实现有效方向迁移,适合跨模型分析
不同家族的大语言模型使用不同的隐藏维度、分词器和训练流程,导致行为方向难以比较或迁移。本文提出锚点投影框架,将各模型的隐藏表示映射至共享锚点坐标空间(ACS)。从源模型提取的行为方向被投影至ACS并平均为标准方向。对于新模型,仅通过锚点激活即可重建该方向,无需微调或目标特定方向提取。我们在五个指令微调模型家族和十个行为轴上进行评估,发现Llama-Qwen-Mistral-Phi(LQMP)集群在ACS中的同轴方向高度对齐。该共享结构可迁移至下游任务:在未见目标上达到0.83的十分类检测准确率和0.95的平均二分类AUROC;标准引导在分布偏移下可引发高达+0.46%的拒绝率变化。敏感性分析表明,仅需两个源模型和小规模锚点池即可近似可迁移方向。总体而言,ACS为跨家族可解释性提供了新视角,揭示表示层面的迁移在模型家族间依然稳健。
原文摘要 · Abstract (English)
Large language models from different families use different hidden dimensions, tokenizers, and training procedures, making behavioral directions difficult to compare or transfer across models. We introduce an anchor-projection framework that maps hidden representations from each model into a shared anchor coordinate space (ACS). Behavioral directions extracted from source models are projected into ACS and averaged into a canonical direction. For a new model, the canonical direction is reconstructed into its native hidden space using only anchor activations, without fine-tuning or target-specific direction extraction. We evaluate five instruction-tuned model families and ten behavioral axes. We find that same-axis directions align tightly across the Llama-Qwen-Mistral-Phi (LQMP) cluster in ACS. This shared structure transfers to downstream tasks. For the aligned LQMP cluster, held-out targets achieve (0.83) ten-way detection accuracy and (0.95) mean binary AUROC, while canonical steering induces refusal-rate shifts of up to +0.46% under distribution shift. Sensitivity analyses show that two source models and small anchor pools already suffice to approximate transferable directions. Overall, ACS provides a novel perspective on cross-family interpretability, revealing that representation-level transfer remains robust across model families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。