通过人格引导提升大模型在多智能体中的协作能力
Identifying Cooperative Personalities in Multi-agent Contexts through Personality Steering with Representation Engineering
- 用表示工程操控大五人格特质,调节模型行为
- 宜人性和尽责性越高,合作率提升但易被利用
- 适合研究AI协作机制与对齐策略的学者
随着大语言模型(LLMs)自主能力增强,其在多智能体环境中的协作变得愈发重要。然而,它们常因缺乏合作而产生次优结果。受阿克塞尔罗德重复囚徒困境(IPD)锦标赛启发,我们探究人格特质如何影响LLM的协作行为。通过表示工程,我们调控大五人格特质(如宜人性、尽责性),并分析其对IPD决策的影响。结果表明,宜人性和尽责性越高,合作表现越好,但同时也更易被对方利用,揭示了基于人格引导在对齐AI智能体方面的潜力与局限。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) gain autonomous capabilities, their coordination in multi-agent settings becomes increasingly important. However, they often struggle with cooperation, leading to suboptimal outcomes. Inspired by Axelrod's Iterated Prisoner's Dilemma (IPD) tournaments, we explore how personality traits influence LLM cooperation. Using representation engineering, we steer Big Five traits (e.g., Agreeableness, Conscientiousness) in LLMs and analyze their impact on IPD decision-making. Our results show that higher Agreeableness and Conscientiousness improve cooperation but increase susceptibility to exploitation, highlighting both the potential and limitations of personality-based steering for aligning AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。