通过连续激活调节,揭示大模型认知行为的动态变化过程。
CBMAS: Cognitive Behavioral Modeling via Activation Steering
- 构建可调节数值路径,实现对模型行为的精细控制。
- 发现微小干预即可引发行为突变的关键转折点。
- 适合关注模型可解释性与行为调控的研究者使用。
大型语言模型在不同提示、层和上下文中常表现出不可预测的认知行为,难以诊断和控制。我们提出CBMAS,一种基于连续激活调节的诊断框架,将认知偏差分析从离散的前后干预扩展为可解释的动态轨迹。通过结合调节向量构造、密集α扫描、基于logit lens的偏差曲线及层-位置敏感性分析,该方法能够揭示小干预强度即引发行为翻转的临界点,并展示调节效应随层深度的变化规律。我们认为,这种连续诊断机制连接了高层行为评估与底层表征动态,推动了大模型的认知可解释性研究。项目仓库提供CLI工具及多种认知行为数据集:https://github.com/shimamooo/CBMAS。
原文摘要 · Abstract (English)
Large language models (LLMs) often encode cognitive behaviors unpredictably across prompts, layers, and contexts, making them difficult to diagnose and control. We present CBMAS, a diagnostic framework for continuous activation steering, which extends cognitive bias analysis from discrete before/after interventions to interpretable trajectories. By combining steering vector construction with dense α-sweeps, logit lens-based bias curves, and layer-site sensitivity analysis, our approach can reveal tipping points where small intervention strengths flip model behavior and show how steering effects evolve across layer depth. We argue that these continuous diagnostics offer a bridge between high-level behavioral evaluation and low-level representational dynamics, contributing to the cognitive interpretability of LLMs. Lastly, we provide a CLI and datasets for various cognitive behaviors at the project repository, https://github.com/shimamooo/CBMAS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。