研究形容词如何影响大模型行为,发现关键形容词有强大但非通用的操控力。
Investigating Linguistic Steering: An Analysis of Adjectival Effects Across Large Language Model Architectures

- 用谢尔普利值量化形容词对模型性能的操控效果,实现可解释的归因分析。
- 少数形容词能显著改变模型输出,其效果在不同模型间差异大且不一致。
- 大模型中形容词间存在复杂交互,小模型则更依赖字面理解,适合不同对齐策略。
实现大语言模型(LLM)的可靠控制需要精确、可扩展地理解其对语言线索的解读方式。本文引入一种基于谢尔普利值的严格框架,量化单个形容词对模型性能的操控效应,超越经验性直觉,实现原理性归因。在涵盖o3、gpt-4o-mini、phi-3、llama-3-70b和deepseek-r1等多样模型的MMLU基准上,测试了100个形容词。结果揭示若干关键发现:首先,少数形容词作为不成比例的强大“杠杆”,但其效应并非普遍适用;跨模型分析显示存在“家族效应”——同源模型表现出相关敏感性特征,而架构迥异的模型反应则高度不相关,挑战了通用提示策略的可行性。其次,深入研究发现,这些强效形容词的作用方向并非固有,而是高度依赖其句法角色与提示中的位置。对于gpt-4o-mini等大模型,首次提供定量证据表明存在强非加性交互效应:形容词可协同增强、相互抑制甚至反转彼此影响;而phi-3等小模型则表现出更字面、较少组合性的响应。结果表明,随着模型规模扩大,其对提示的理解变得更复杂但也更不可预测,这对稳健控制模型行为构成重大挑战,凸显了组合式与模型特异性对齐技术的必要性。
原文摘要 · Abstract (English)
Achieving reliable control of Large Language Models (LLMs) requires a precise, scalable understanding of how they interpret linguistic cues. We introduce a rigorous framework using Shapley values to quantify the steering effect of individual adjectives on model performance, moving beyond anecdotal heuristics to principled attribution. Applying this method to 100 adjectives across a diverse suite of models (including o3, gpt-4o-mini, phi-3, llama-3-70b, and deepseek-r1) on the MMLU benchmark, we uncover several critical findings for AI alignment. First, we find that a small subset of adjectives act as disproportionately powerful "levers," yet their effects are not universal. Cross-model analysis reveals a "family effect": models of a shared lineage exhibit correlated sensitivity profiles, while architecturally distinct models react in a largely uncorrelated manner, challenging the notion of a one-size-fits-all prompting strategy. Second, focused follow-up studies demonstrate that the steering direction of these powerful adjectives is not intrinsic but is highly contingent on their syntactic role and position within the prompt. For larger models like gpt-4o-mini, we provide the first quantitative evidence of strong, non-additive interaction effects where adjectives can synergistically amplify, antagonistically dampen, or even reverse each other's impact. In contrast, smaller models like phi-3 exhibit a more literal and less compositional response. These results suggest that as models scale, their interpretation of prompts becomes more sophisticated but also less predictable, posing a significant challenge for robustly steering model behavior and highlighting the need for compositional and model-specific alignment techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。