发现语言模型内部存在可操控的社会角色粒度轴。
The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models

- 通过对比宏观与微观角色隐藏状态,构建粒度轴。
- 该轴解释52.6%方差,角色投影随粒度单调上升。
- 可跨模型操控响应粒度,适合可控角色生成研究。
大型语言模型(LLMs)常被提示扮演从个体到机构等社会角色,但其内部表征是否编码了角色粒度(从个人经验到组织、机构或国家层面推理)尚不明确。本文证明其确实存在。定义基于对比的粒度轴为宏观与微观角色隐藏状态的均值差。在Qwen3-8B中,该轴与角色表示空间主成分(PC1)余弦相似度达0.972,解释52.6%方差,表明粒度是组织提示角色的核心几何轴。构建涵盖五级粒度的75个社会角色,收集91,200条条件响应,提取角色级隐藏状态并投影至轴上。投影值在所有层级单调递增,跨层、提示变体、端点定义、留出集及评分过滤子集均稳定,并可迁移至Llama-3.1-8B-Instruct。轴具因果相关性:沿轴操纵激活可按预期改变响应粒度,如在允许本地回应的提示下,Llama模型粒度评分由2.00升至3.17。两模型控制能力差异,表明操纵效果取决于模型默认运行模式。整体表明,社会角色粒度不仅是风格表面特征,更是结构化、有序且可因果操控的潜在方向。
原文摘要 · Abstract (English)
Large language models (LLMs) are routinely prompted to take on social roles ranging from individuals to institutions, yet it remains unclear whether their internal representations encode the granularity of such roles, from micro-level individual experience to macro-level organizational, institutional, or national reasoning. We show that they do. We define a contrast-based Granularity Axis as the difference between mean macro- and micro-role hidden states. In Qwen3-8B, this axis aligns with the principal axis (PC1) of the role representation space at cosine 0.972 and accounts for 52.6% of its variance, indicating that granularity is the dominant geometric axis organizing prompted social roles. We construct 75 social roles across five granularity levels and collect 91,200 role-conditioned responses over shared questions and prompt variants, then extract role-level hidden states and project them onto the axis. Role projections increase monotonically across all five levels, remain stable across layers, prompt variants, endpoint definitions, held-out splits, and score-filtered subsets, and transfer to Llama-3.1-8B-Instruct. The axis is also causally relevant: activation steering along it shifts response granularity in the predicted direction, with Llama moving from 2.00 to 3.17 on a five-point macro scale under positive steering on prompts that admit local responses. The two models differ in controllability, suggesting that steering depends on each model's default operating regime. Overall, our findings suggest that social role granularity is not merely a stylistic surface feature, but a structured, ordered, and causally manipulable latent direction in role-conditioned language model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。