arXiv:2604.02145cs.AIcs.CL2026-04被引 2

为大模型设计行为画像系统,揭示其性格特质与能力无关。

MTI: A Behavior-Based Temperament Profiling System for AI Agents

  • 基于行为测试构建四维性格量表,不依赖自我描述。
  • 发现性格维度独立,小模型间性格差异显著但与大小无关。
  • 适合研究模型行为差异或评估对齐效果的团队使用。

同等能力的大模型可能表现出根本不同的行为模式,但目前尚无标准化工具衡量这些气质差异。现有方法要么借用人类人格维度并依赖自评(与大模型实际行为不符),要么将行为变异视为缺陷而非特质。我们提出模型气质指数(MTI),一种基于行为的代理气质分析系统,从四个维度测量:反应性(环境敏感度)、合规性(指令-行为一致性)、社会性(关系资源分配)和韧性(抗压能力)。依托模型医学中的四壳模型,MTI通过结构化测试协议,采用两阶段设计,将能力与倾向分离。我们对10个小型语言模型(1.7B-9B参数,6家机构,3种训练范式)进行画像,报告五大发现:(1)指令微调模型中四维度基本独立(所有|r| < 0.42);(2)维度内特征可分离——合规性分为独立的正式与立场子维度(r=0.002),韧性则包含反向关联的认知与对抗子维度;(3)合规性与韧性存在悖论,观点输出与事实脆弱性通过独立通道运作;(4)强化学习人类反馈(RLHF)不仅改变轴向得分,还创造基线模型中不存在的内部特征分化;(5)气质与模型规模(1.7B-9B)无关,确认MTI测量的是倾向而非能力。

原文摘要 · Abstract (English)

AI models of equivalent capability can exhibit fundamentally different behavioral patterns, yet no standardized instrument exists to measure these dispositional differences. Existing approaches either borrow human personality dimensions and rely on self-report (which diverges from actual behavior in LLMs) or treat behavioral variation as a defect rather than a trait. We introduce the Model Temperament Index (MTI), a behavior-based profiling system that measures AI agent temperament across four axes: Reactivity (environmental sensitivity), Compliance (instruction-behavior alignment), Sociality (relational resource allocation), and Resilience (stress resistance). Grounded in the Four Shell Model from Model Medicine, MTI measures what agents do, not what they say about themselves, using structured examination protocols with a two-stage design that separates capability from disposition. We profile 10 small language models (1.7B-9B parameters, 6 organizations, 3 training paradigms) and report five principal findings: (1) the four axes are largely independent among instruction-tuned models (all |r| < 0.42); (2) within-axis facet dissociations are empirically confirmed -- Compliance decomposes into fully independent formal and stance facets (r = 0.002), while Resilience decomposes into inversely related cognitive and adversarial facets; (3) a Compliance-Resilience paradox reveals that opinion-yielding and fact-vulnerability operate through independent channels; (4) RLHF reshapes temperament not only by shifting axis scores but by creating within-axis facet differentiation absent in the unaligned base model; and (5) temperament is independent of model size (1.7B-9B), confirming that MTI measures disposition rather than capability.

大模型行为性格分析评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。