用五大性格特质评估大模型行为,揭示不同操控方法的优劣与代价。
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
- 对比三种性格操控方法:上下文学习、参数高效微调和机制引导。
- 发现上下文学习对能力影响最小,微调效果最好但性能下降明显。
- 提出性格净化与稳定性框架,适合部署优化与可解释性研究者参考。
大语言模型中的性格操控在客服和智能体场景中日益普及,但其机制与权衡尚不清晰。本文基于五大性格特质,系统比较了上下文学习(ICL)、参数高效微调(PEFT)和机制引导(MS)三种方法。构建了平衡高低特质响应的对比数据集,实现有效引导向量计算与公平评估;提出基于运行内Δ分析的统一评估框架,分离推理能力、代理表现与人口统计偏差,覆盖MMLU、GAIA和BBQ基准;开发性格净化技术以区分开放性与尽责性,缓解特征编码重叠问题;提出三级稳定性框架,量化方法、特质及组合层面的鲁棒性。在Gemma-2-2B-IT与LLaMA-3-8B-Instruct上的实验表明:ICL对能力影响最小,性能损失少;PEFT达到最高对齐度但任务性能下降显著;MS提供轻量级实时控制,效果接近微调。性格层面分析显示,开放性最难操控,宜人性最抵抗ICL,性格编码集中在中间层。结果表明,性格操控是多层次行为表征探针,连接表面调控、参数编码与激活层级引导,机制引导为部署与可解释性提供轻量替代方案。
原文摘要 · Abstract (English)
Personality manipulation in large language models (LLMs) is increasingly applied in customer service and agentic scenarios, yet its mechanisms and trade-offs remain unclear. We present a systematic study of personality control using the Big Five traits, comparing in-context learning (ICL), parameter-efficient fine-tuning (PEFT), and mechanistic steering (MS). Our contributions are fourfold. First, we construct a contrastive dataset with balanced high/low trait responses, enabling effective steering vector computation and fair cross-method evaluation. Second, we introduce a unified evaluation framework based on within-run $Δ$ analysis that disentangles, reasoning capability, agent performance, and demographic bias across MMLU, GAIA, and BBQ benchmarks. Third, we develop trait purification techniques to separate openness from conscientiousness, addressing representational overlap in trait encoding. Fourth, we propose a three-level stability framework that quantifies method-, trait-, and combination-level robustness, offering practical guidance under deployment constraints. Experiments on Gemma-2-2B-IT and LLaMA-3-8B-Instruct reveal clear trade-offs: ICL achieves strong alignment with minimal capability loss, PEFT delivers the highest alignment at the cost of degraded task performance, and MS provides lightweight runtime control with competitive effectiveness. Trait-level analysis shows openness as uniquely challenging, agreeableness as most resistant to ICL, and personality encoding consolidating around intermediate layers. Taken together, these results establish personality manipulation as a multi-level probe into behavioral representation, linking surface conditioning, parameter encoding, and activation-level steering, and positioning mechanistic steering as a lightweight alternative to fine-tuning for both deployment and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。