语言结构变化会显著影响大模型立场判断,且关键影响部位可定位。
Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing

- 通过控制句法重构实验,验证不同表达方式改变模型立场
- 中晚期解码层在修复立场偏差时效果最强,尤其最后一句位置的块输出
- 适合研究模型决策机制、提示工程与偏见控制的研究者
大型语言模型对提示和输入形式敏感,但现有研究多关注词汇层面,忽视句法结构选择的影响。本文以政治立场判断为语义敏感任务,扩展英文政治陈述数据集,生成六种受控句法重写类型,保持或反转原意。在四个开源模型上实验发现,无论语义是否保留,句法重构均引发立场不稳定性。为进一步定位影响节点,采用激活修补技术:将原句激活值替换到重写句的前向传播中,测量哪些组件能恢复原始立场分布。结果显示,中晚期解码层,特别是最终提示位置的块输出,提供最强的立场恢复信号。
原文摘要 · Abstract (English)
Large language models (LLMs) are known to be sensitive to prompt and input formulations. However, existing studies have focused on lexical realization and largely ignored constructional choice. This paper studies whether linguistic construction can systematically shift LLM decisions and where these shifts can be causally localized inside the model. We use political stance judgment as a meaning-sensitive case study and extend an English political statements dataset, resulting in six controlled linguistic rewrite types that preserve or invert the meaning of a statement. Experiments on four open-weight models show that stance instability affect both meaning-preserving and meaning-inversing rewrites. Because output shifts reveal that rewrites affect stance, but not where in the model, we apply activation patching, where activations from the original statement are substituted into the forward pass for the rewritten statement and measure which components recover the original stance distribution. The results show that mid-to-late decoder layers, especially block outputs at the final prompt position, provide the strongest restoration signal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。