通过可解释的特征操控,让大模型角色扮演更稳定有效
Improving LLM Reasoning through Interpretable Role-Playing Steering
- 用稀疏自编码器识别角色扮演的内部特征并生成可控引导向量
- Llama3.1-8B在CSQA上零样本推理准确率提升至39.80%
- 相比传统提示法,结果更稳定且能揭示模型内部机制
角色扮演已成为提升大语言模型(LLMs)推理能力的有效方法。然而,现有方法主要依赖提示工程,常缺乏稳定性与可解释性。本文提出稀疏自编码器角色扮演引导(SRPS)框架,通过提取角色提示的潜在表征,基于激活模式选择最相关特征,并构建可注入残差流的引导向量,实现对角色行为的细粒度控制。该方法还能揭示角色信息如何影响模型内部激活。在多个推理基准和模型规模上进行的大量实验表明,性能持续提升:在零样本链式思维(CoT)设置下,Llama3.1-8B在CSQA上的准确率从31.86%提升至39.80%,Gemma2-9B在SVAMP上的准确率从37.50%提升至45.10%。结果表明,SRPS能有效增强LLM推理能力,相较传统提示法具有更高的可解释性与稳定性。
原文摘要 · Abstract (English)
Role-playing has emerged as an effective technique for enhancing the reasoning capabilities of large language models (LLMs). However, existing methods primarily rely on prompt engineering, which often lacks stability and interpretability. In this paper, we introduce Sparse Autoencoder Role-Playing Steering (SRPS), a novel framework that identifies and manipulates internal model features associated with role-playing behavior. Our approach extracts latent representations from role-play prompts, selects the most relevant features based on activation patterns, and constructs a steering vector that can be injected into the model's residual stream with controllable intensity. Our method enables fine-grained control over role-specific behavior and offers insights into how role information influences internal model activations. Extensive experiments across various reasoning benchmarks and model sizes demonstrate consistent performance gains. Notably, in the zero-shot chain-of-thought (CoT) setting, the accuracy of Llama3.1-8B on CSQA improves from 31.86% to 39.80%, while Gemma2-9B on SVAMP increases from 37.50% to 45.10%. These results highlight the potential of SRPS to enhance reasoning ability in LLMs, providing better interpretability and stability compared to traditional prompt-based role-playing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。