arXiv:2510.03799cs.CLcs.AI2025-10被引 2

发现大模型能识别并生成政治认知框架。

Mechanistic Interpretability of Socio-Political Frames in Language Models

  • 通过隐藏层分析定位关键认知框架
  • 零样本下准确识别'严格父亲'与'养育父母'框架
  • 为理解模型如何表达人类概念提供新视角

本文探讨大语言模型在社会政治语境中生成和识别深层认知框架的能力。我们证明,大型语言模型在生成体现特定框架的文本方面高度流畅,并能在零样本设置下识别这些框架。受机制可解释性研究启发,我们探究了'严格父亲'和'养育父母'框架在模型隐藏表示中的位置,发现与它们存在强相关性的单一维度。研究结果有助于理解大语言模型如何捕捉和表达有意义的人类概念。

原文摘要 · Abstract (English)

This paper explores the ability of large language models to generate and recognize deep cognitive frames, particularly in socio-political contexts. We demonstrate that LLMs are highly fluent in generating texts that evoke specific frames and can recognize these frames in zero-shot settings. Inspired by mechanistic interpretability research, we investigate the location of the `strict father' and `nurturing parent' frames within the model's hidden representation, identifying singular dimensions that correlate strongly with their presence. Our findings contribute to understanding how LLMs capture and express meaningful human concepts.

认知框架可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。