arXiv:2608.05164cs.CLcs.LG2026-08被引 1

大模型间共享语义表示,可实现无微调的跨模型行为控制。

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

论文配图:Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study
图 1 · 摘自论文原文
  • 基于不同架构的大模型存在几何相似的语义方向,可跨模型传递控制向量。
  • 参数量≥1.7B时,47%~49%的跨模型特征对有效(相关系数≥0.60)。
  • 仅需一个通用向量,即可在4/5模型中实现67.3%的控制准确率。

独立训练的大语言模型虽架构不同,但可能发展出共享的语义概念内部表征——这种几何相似性是否具备跨模型行为控制的功能意义尚不清楚。本文首次系统评估了跨模型引导迁移,发现共享的模型几何结构在特定条件下可被功能利用:当模型具备足够表征容量时,一个模型的语义方向可控制另一独立训练模型。研究涵盖五个开源模型(参数量0.8B–8B,两种架构谱系),每个模型训练一个稀疏自编码器覆盖15个语义领域,测试所有20对模型间的定向对齐。在约1.7B参数处观察到显著断点:≥1.7B时,47%–49%的跨模型特征对通过验证(皮尔逊相关系数≥0.60,Procrustes余弦0.895–0.956),低于0.8B则急剧下降。跨模型引导向量(B3-TI)在15个受监督概念上取得71.0%胜率,优于同模型原生向量的68.0%;单一通用向量在5个模型中的4个实现67.3%准确率,无需每模型单独监督。性能随模型规模<1.7B或生成不稳定性而下降,证实功能可用性依赖于充分的表征能力。研究强调了机制可解释性中的规模阈值重要性:在7B规模验证的工具未必适用于小模型,须重新验证。本研究为柏拉图表征假说提供了首个功能性补充——独立训练模型间几何收敛支持无需微调的跨模型行为控制,前提为满足特定规模条件。

原文摘要 · Abstract (English)

Independently trained large language models may develop shared internal representations of semantic concepts despite architectural differences -- but whether this geometric similarity has functional consequences for cross-model behavioural control remains untested. We present the first systematic evaluation of cross-model steering transfer and show that shared LLM geometry is functionally exploitable, conditionally: concept directions from one model can steer a different independently trained model when sufficient representational capacity exists. We study five open-weight models spanning three parameter scales (0.8B--8B) and two architectural lineages, training one Sparse Autoencoder per model across 15 semantic domains and testing alignment across all 20 directed model pairs. We observe a suggestive discontinuity near 1.7B parameters: at >= 1.7B scale, 47--49% of cross-model feature pairs validate (Pearson r >= 0.60, Procrustes cosines 0.895--0.956), while alignment degrades sharply below 0.8B. Cross-model steering vectors (B3-TI) achieve a 71.0% win rate across 15 supervised concepts versus 68.0% for same-model native vectors; a single universal vector achieves 67.3% in 4 of 5 models without any per-model supervision. Transfer degrades for models below 1.7B and for one model with generation instability, confirming that functional exploitability requires sufficient representational capacity. Our findings underscore the importance of scale thresholds in mechanistic interpretability: tools validated at 7B scale may not transfer to smaller models without revalidation. We provide the first functional complement to the Platonic Representation Hypothesis -- geometric convergence across independently trained LLMs supports cross-model behavioural control without fine-tuning, under the identified scale conditions.

大模型语义控制跨模型迁移可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。