通过真实交互场景数据,让大模型行为风格可测量、可控制。
Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

- 构建3200个对比行为场景,基于心理学量表分析模型行为模式。
- 发现不同模型有稳定的行为特征,且在不同任务中表现各异。
- 提出行为模式轴,用思考过程控制行为风格更精准可靠。
大型语言模型在交互场景中的行为风格影响用户体验、安全性和决策效果。现有研究多依赖第一人称自我报告问卷,易受表达方式影响,缺乏对具体行为的实证基础。本文提出一种基于真实交互情境的行为数据(B-data)框架,构建了3200个涵盖20种行为模式和4种提示模态的对比场景,依据BFI-2、DOSPERT、HEXACO等经验证的心理维度设计。实验发现,大模型表现出稳定且模型特异的行为画像,并在第一人称决策、建议提供与任务执行中呈现模态依赖性变化。进一步提出行为模式轴(BMAs),通过对比行为轨迹提取激活空间方向实现行为控制。相比基于输出响应的BMAs,基于思考过程的BMAs更准确反映意图机制,控制效果更清晰。结果表明,大模型的类人格倾向应理解为可测量、可调控的具身化行为模式,而非抽象自述特质。代码与数据已开源。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report questionnaires administered in first-person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in concrete model behavior. In this work, we introduce a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality. We construct 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric facets such as BFI-2, DOSPERT, and HEXACO. Using this framework, we find that LLMs exhibit stable and model-specific behavioral profiles, while also revealing register-dependent shifts across first-person decisions, advice-giving, and task execution. We then show that these behavioral patterns can be controlled through Behavioral Mode Axes (BMAs), activation-space directions derived from contrastive behavioral traces. Compared with response-derived BMAs, which are more prone to trait drift, thought-derived BMAs more faithfully capture the intended behavioral mechanism and provide cleaner control over situated behavioral styles. Our results suggest that LLM personality-like tendencies are better understood not as abstract self-report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts. Our code and data are available at https://github.com/lhz191/LLM-Behavioral-Personality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。