研究大模型如何像人一样形成对提示的刻板印象。
Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
- 用线性探测分析模型隐藏层中的印象特征。
- 模型对提示的印象不一致,但隐藏层可稳定解码。
- 提示的风格和方言特征影响模型生成的印象。
我们引入并研究了人工印象——大语言模型在内部表示中对提示形成的、类似人类基于语言的刻板印象与偏见的模式。通过在生成的提示上拟合线性探测器,预测其在二维刻板印象内容模型(SCM)下的印象。利用这些探测器,我们考察了印象与下游模型行为的关系,以及可能影响印象的提示特征。结果发现,当被提示时,模型对印象的报告并不一致;但其隐藏表示中的人工印象却更一致地可线性解码。此外,我们证明提示的人工印象能预测模型回复中模糊表达(hedging)的质量与使用程度。我们还研究了提示中特定内容、风格和方言特征如何影响模型产生的印象。
原文摘要 · Abstract (English)
We introduce and study artificial impressions--patterns in LLMs' internal representations of prompts that resemble human impressions and stereotypes based on language. We fit linear probes on generated prompts to predict impressions according to the two-dimensional Stereotype Content Model (SCM). Using these probes, we study the relationship between impressions and downstream model behavior as well as prompt features that may inform such impressions. We find that LLMs inconsistently report impressions when prompted, but also that impressions are more consistently linearly decodable from their hidden representations. Additionally, we show that artificial impressions of prompts are predictive of the quality and use of hedging in model responses. We also investigate how particular content, stylistic, and dialectal features in prompts impact LLM impressions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。