将大模型谄媚行为拆解为心理特质组合,实现可解释干预
Sycophancy as compositions of Atomic Psychometric Traits
- 把谄媚行为看作情绪、开放性等心理特质的几何组合
- 通过激活向量操作发现高外向+低尽责易诱发谄媚
- 支持加减投影等可解释干预,适合安全可控模型设计
谄媚是大模型的关键行为风险,常被视作单一因果机制导致的孤立故障。本文提出将其建模为情绪性、开放性、宜人性等心理特质的几何与因果组合,类似心理测量学中的因子分解。利用对比激活添加(CAA)方法,将激活方向映射至这些因素,并研究不同组合如何引发谄媚行为(如高外向性结合低尽责性)。该视角支持可解释且可组合的向量级干预,如加法、减法和投影,可用于缓解大模型中的安全关键行为。
原文摘要 · Abstract (English)
Sycophancy is a key behavioral risk in LLMs, yet is often treated as an isolated failure mode that occurs via a single causal mechanism. We instead propose modeling it as geometric and causal compositions of psychometric traits such as emotionality, openness, and agreeableness - similar to factor decomposition in psychometrics. Using Contrastive Activation Addition (CAA), we map activation directions to these factors and study how different combinations may give rise to sycophancy (e.g., high extraversion combined with low conscientiousness). This perspective allows for interpretable and compositional vector-based interventions like addition, subtraction and projection; that may be used to mitigate safety-critical behaviors in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。