让大模型根据情境动态调整性格,更自然地与人互动。
Beyond Static Personas: Situational Personality Steering for Large Language Models

- 通过分析性格神经元,发现性格会随情境变化
- 无需训练,用神经元检索与加权实现情境化性格控制
- 在多个测试集上表现优于现有方法,适配不同模型
个性化大语言模型(LLMs)能提升人机交互的自然度。然而,现有方法受限于可控性差、资源消耗高,且依赖静态性格建模,难以适应不同情境。我们通过多角度分析性格神经元,首次验证了性格的情境依赖性及稳定的行为模式。基于此,提出IRIS——一种无需训练、基于神经元的“识别-检索-调节”框架,实现情境化性格调控。该方法包括情境性格神经元识别、情境感知神经元检索与相似性加权调节。我们在PersonalityBench和新构建的SPBench(情境性格基准)上进行实证验证,结果表明,该方法在复杂未见情境下仍具备良好泛化与鲁棒性,超越当前最优基线,适用于多种模型架构。
原文摘要 · Abstract (English)
Personalized Large Language Models (LLMs) facilitate more natural, human-like interactions in human-centric applications. However, existing personalization methods are constrained by limited controllability and high resource demands. Furthermore, their reliance on static personality modeling restricts adaptability across varying situations. To address these limitations, we first demonstrate the existence of situation-dependency and consistent situation-behavior patterns within LLM personalities through a multi-perspective analysis of persona neurons. Building on these insights, we propose IRIS, a training-free, neuron-based Identify-Retrieve-Steer framework for advanced situational personality steering. Our approach comprises situational persona neuron identification, situation-aware neuron retrieval, and similarity-weighted steering. We empirically validate our framework on PersonalityBench and our newly introduced SPBench, a comprehensive situational personality benchmark. Experimental results show that our method surpasses best-performing baselines, demonstrating IRIS's generalization and robustness to complex, unseen situations and different models architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。