发现大模型有独特行为指纹,非随机反应偏差。
Machine individuality: Separating genuine idiosyncrasy from response bias in large language models
- 用交叉随机效应模型分离出模型真实个性差异。
- 16.9%的差异源于刺激特定的个体性,非噪声或偏见。
- 每款模型都有独特行为指纹,适合评估模型稳定性。
随着大语言模型(LLMs)广泛应用于决策支持与陪伴等场景,理解其行为特征变得至关重要。现有方法通过心理量表和认知范式刻画模型倾向,但无法区分行为差异是稳定的、针对特定刺激的个体性,还是全局响应偏见与随机噪声所致。本研究采用心理测量学中常用的交叉随机效应模型,分析10个开源权重大模型对超过10万词、14个心理语言学规范的7490万条评分。结果显示,平均16.9%的方差可归因于刺激特定的个体性,显著高于统计零模型。跨规范预测分析表明,这种个体性构成一致的行为指纹,每款模型各不相同。这些差异无法归因于响应偏见或随机噪声,我们称之为机器个体性。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly integrated into daily life, in roles ranging from high-stakes decision support to companionship, understanding their behavioral dispositions becomes critical. A growing literature uses psychometric inventories and cognitive paradigms to profile LLM dispositions. However, these approaches cannot determine whether behavioral differences reflect stable, stimulus-specific individuality or global response biases and stochastic noise. Here, we apply crossed random-effects models -- widely used in psychometrics to separate systematic effects -- to 74.9 million ratings provided by 10 open-weight LLMs for over 100,000 words across 14 psycholinguistic norms. On average, 16.9% of variance is attributable to stimulus-specific individuality, robustly exceeding a statistical null model. Cross-norm prediction analyses reveal this individuality as a coherent fingerprint, unique to each model. These results identify individual differences among LLMs that cannot be attributed to response biases or stochastic noise. We term these differences machine individuality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。