给大模型的个性行为建模,用客观方式描述其可信、友善等特质。
Evaluating Language Model Character Traits
- 从行为一致性出发,定义大模型的可信、讨好等性格特征
- 模型规模、微调和提示方式影响性格表现的一致性
- 适合关注模型行为稳定性与可解释性的研究人员
语言模型(LMs)可能表现出类人行为,但如何描述这些行为而不过度拟人化仍不明确。本文提出一种行为主义视角下的语言模型性格特质形式化框架:如诚实性、讨好性或连贯信念与意图等,表现为一致的行为模式。该理论基于实证发现,即语言模型展现出准确且逻辑一致的信念,以及有益无害的意图。研究发现,模型在特定情境下展现的性格特质(如诚实性和危害性)具有稳定性,但在其他情境中则可能呈现反射性,即反映前序交互中的行为。性格特质的一致性受模型规模、微调和提示方式的影响。该形式化方法使我们能以直观但非拟人化的方式精准描述模型行为。
原文摘要 · Abstract (English)
Language models (LMs) can exhibit human-like behaviour, but it is unclear how to describe this behaviour without undue anthropomorphism. We formalise a behaviourist view of LM character traits: qualities such as truthfulness, sycophancy, or coherent beliefs and intentions, which may manifest as consistent patterns of behaviour. Our theory is grounded in empirical demonstrations of LMs exhibiting different character traits, such as accurate and logically coherent beliefs, and helpful and harmless intentions. We find that the consistency with which LMs exhibit certain character traits varies with model size, fine-tuning, and prompting. In addition to characterising LM character traits, we evaluate how these traits develop over the course of an interaction. We find that traits such as truthfulness and harmfulness can be stationary, i.e., consistent over an interaction, in certain contexts, but may be reflective in different contexts, meaning they mirror the LM's behavior in the preceding interaction. Our formalism enables us to describe LM behaviour precisely in intuitive language, without undue anthropomorphism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。