arXiv:2509.16332cs.AIcs.CL2025-09被引 3

通过人格特质调节,可显著影响大模型的能力与安全性。

Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models

  • 基于五大性格模型调控人格特质,观察其对模型行为的影响。
  • 降低尽责性导致安全指标(如WMDP、TruthfulQA)和通用能力(MMLU)下降。
  • 为模型安全评估与部署后行为控制提供新思路,适合关注模型可控性的研究者。

大语言模型日益参与高风险交互,推动对其能力与安全性的研究。尽管已有研究表明大模型具备稳定可测的合成人格特征,但人格特质如何影响模型行为仍不明确。本文基于五大性格框架,探究人格调控对模型在能力与安全基准上的影响。实验显示:降低尽责性会显著削弱模型在WMDP、TruthfulQA、ETHICS、Sycophancy等安全相关基准的表现,同时导致通用能力(以MMLU衡量)下降。结果表明,人格塑造是影响模型安全与综合能力的重要且未被充分探索的控制维度。本文讨论了对安全评估、对齐策略、部署后行为引导的意义,并警示可能被滥用的风险。研究呼吁开展人格敏感的安全评估与动态行为控制的新方向。

原文摘要 · Abstract (English)

Large Language Models increasingly mediate high-stakes interactions, intensifying research on their capabilities and safety. While recent work has shown that LLMs exhibit consistent and measurable synthetic personality traits, little is known about how modulating these traits affects model behavior. We address this gap by investigating how psychometric personality control grounded in the Big Five framework influences AI behavior in the context of capability and safety benchmarks. Our experiments reveal striking effects: for example, reducing conscientiousness leads to significant drops in safety-relevant metrics on benchmarks such as WMDP, TruthfulQA, ETHICS, and Sycophancy as well as reduction in general capabilities as measured by MMLU. These findings highlight personality shaping as a powerful and underexplored axis of model control that interacts with both safety and general competence. We discuss the implications for safety evaluation, alignment strategies, steering model behavior after deployment, and risks associated with possible exploitation of these findings. Our findings motivate a new line of research on personality-sensitive safety evaluations and dynamic behavioral control in LLMs.

人格建模模型安全行为控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。