arXiv:2504.08399cs.CLcs.AI2025-04EMNLP被引 5

用多个角色观察者评估大模型人格,比自我报告更准。

Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models

  • 让不同关系角色的代理与大模型对话后评分。
  • 5到7个观察者评分最可靠,偏差最小。
  • 适合研究模型人格评估或心理测评的学者。

传统自评问卷评估大语言模型人格时,因偏见和元知识污染难以捕捉行为细节。本文提出一种基于心理学信息源报告方法的多观察者框架:不依赖自我评价,而是配置具有特定关系背景(如家人、朋友、同事)的多个观察者代理,与目标模型对话后,从大五人格维度评估其行为。实验表明,观察者评分与人类判断更一致,且揭示了大模型自评中的系统性偏差。当观察者数量为5至7人时,偏差最小,可靠性最优。结果表明关系语境影响人格感知,多观察者范式能提供更可靠、情境敏感的模型人格评估方法。

原文摘要 · Abstract (English)

Self-report questionnaires have long been used to assess LLM personality traits, yet they fail to capture behavioral nuances due to biases and meta-knowledge contamination. This paper proposes a novel multi-observer framework for personality trait assessments in LLM agents that draws on informant-report methods in psychology. Instead of relying on self-assessments, we employ multiple observer agents. Each observer is configured with a specific relational context (e.g., family member, friend, or coworker) and engages the subject LLM in dialogue before evaluating its behavior across the Big Five dimensions. We show that these observer-report ratings align more closely with human judgments than traditional self-reports and reveal systematic biases in LLM self-assessments. We also found that aggregating responses from 5 to 7 observers reduces systematic biases and achieves optimal reliability. Our results highlight the role of relationship context in perceiving personality and demonstrate that a multi-observer paradigm offers a more reliable, context-sensitive approach to evaluating LLM personality traits.

人格评估多代理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。