用几何投影矩阵实现更精准的语义控制,避免单一方向的局限。
Conceptors for Semantic Steering

- 用双向概念激活数据估算软投影矩阵,保留多维语义子空间。
- 在三模型三维度上预测概念可分性相关性达r=0.96,无需调参。
- 支持逻辑运算组合,减少无效输出,适合需要可控生成的场景。
基于激活的语义控制可在推理时调节大模型行为,但主流方法将每个概念简化为单一方向,其几何结构未被充分研究。本文提出使用概念器(conceptors):从双边概念激活中聚合估计的软投影矩阵,保留概念的完整多维子空间。几何分析表明,双边子空间严格包含单向量基线。进一步发现,概念器配额可作为无参层选择诊断工具,在三个指令微调模型和三个语义维度上,预测概念可分性的皮尔逊相关系数最高达r=0.96。除层选择外,概念器支持闭式布尔代数(AND、OR、NOT),我们在主题相关子概念上评估了其组合性。在系统五轴设计空间测试中,当概念子空间为多维时,概念器性能优于或相当加法基线,同时显著减少退化输出。概念器控制是一种几何合理、可组合且实际更安全的替代方案,仅需少量对比对即可实现语义引导。
原文摘要 · Abstract (English)
Activation-based steering provides control of LLM behavior at inference time, but the dominant paradigm reduces each concept to a single direction whose geometry is left largely unexamined. Rather than selecting a single steering direction, we use conceptors: soft projection matrices estimated from activations pooled across both poles of a bipolar concept, which preserve the concept's full multidimensional subspace. A geometric analysis shows the bipolar subspace strictly subsumes the single-vector baseline. We further show that the conceptor quota provides a parameter-free layer-selection diagnostic, predicting concept separability with Pearson correlations up to r=0.96 across three instruction-tuned models and three semantic dimensions. Beyond selection, conceptors admit a closed-form Boolean algebra (AND, OR, NOT): we evaluate conceptor compositionality on thematically related sub-concepts. Across a systematic five-axis design-space evaluation, conceptors match or outperform additive baselines at layers where concept subspaces are multi-dimensional while producing substantially fewer degenerate outputs. Conceptor steering is a geometrically principled, compositional, and practically safer alternative to single-direction steering from a limited number of contrastive pairs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。