arXiv:2510.04484cs.CLcs.AI2025-10中稿 · ACL被引 7

评测大模型情感与人格操控的有效性与可信度,发现不同方法各有优劣。

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

  • 提出PsySET基准,评估四种模型在情感与人格上的控制效果。
  • 提示词控制有效但强度难调,向量注入更精细但略降输出质量。
  • 发现情绪操控可能引发安全风险,如快乐降低事实抗性,愤怒提升毒性。

控制大模型模拟的情绪状态和人格特质,是实现社交互动中人性化交互的关键一步。本文提出PsySET——一个心理启发式基准,用于评估不同模型家族在情感与人格领域中的可控性与可信度。研究涵盖四种来自不同大模型家族的模型,结合提示、微调和表征工程等多种操控策略。结果显示,提示法持续有效但难以精确控制强度;而向量注入可实现更精细调节,但略微降低输出质量。此外,我们通过安全性、真实性、公平性和伦理维度评估操控后的可信度,揭示潜在副作用与行为变化:例如,即使是积极情绪‘喜悦’也可能削弱对对抗性事实的鲁棒性、降低隐私意识并加剧偏好偏见;而‘愤怒’虽明显提升毒性,却增强信息泄露抵抗能力。本框架首次实现对情感与人格操控的全面评估,为社交应用中的可解释性与可靠性提供洞见。

原文摘要 · Abstract (English)

The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactions in socially interactive settings. We introduce PsySET, a Psychologically-informed benchmark to evaluate LLM Steering Effectiveness and Trustworthiness across the emotion and personality domains. Our study spans four models from different LLM families paired with various steering strategies, including prompting, fine-tuning, and representation engineering. Our results indicate that prompting is consistently effective but limited in intensity control, whereas vector injections achieve finer controllability while slightly reducing output quality. Moreover, we explore the trustworthiness of steered LLMs by assessing safety, truthfulness, fairness, and ethics, highlighting potential side effects and behavioral shifts. Notably, we observe idiosyncratic effects; for instance, even a positive emotion like joy can degrade robustness to adversarial factuality, lower privacy awareness, and increase preferential bias. Meanwhile, anger predictably elevates toxicity yet strengthens leakage resistance. Our framework establishes the first holistic evaluation of emotion and personality steering, offering insights into its interpretability and reliability for socially interactive applications.

大模型操控情感控制可信度评估人格建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。