arXiv:2603.23485cs.CLcs.AI2026-03被引 2

大模型在相同语法下输出不一致,可能因无关上下文引入偏见。

Failure of contextual invariance in large language models

  • 通过控制代词任务,发现微小语境变化导致模型输出显著偏离。
  • 引入语境后,文化性别刻板印象相关性下降甚至消失。
  • 即使排除重复代词等简单因素,仍有19%~52%案例存在上下文依赖。

标准评估假设大语言模型在语义等价语境中输出稳定。本文在性别推断场景下检验该假设,使用受控的代词选择任务,引入理论上无信息的最小语境,发现模型输出出现大规模系统性偏差。原本在脱离语境时存在的文化性别刻板印象相关性,在引入语境后减弱或消失;而与目标无关的代词性别等理论无关特征,反而成为预测模型行为的最佳指标。基于‘默认上下文性’的分析显示,在跨模型19%至52%的情况下,这种依赖关系在扣除所有边际效应后依然存在,且无法归因于简单的代词重复。结果表明,大模型输出在几乎相同的句法结构下仍违反上下文不变性,对偏见评估与高风险场景部署具有重要影响。

原文摘要 · Abstract (English)

Standard evaluation practices assume that large language model (LLM) outputs are stable when prompts are embedded in contextually equivalent discourses. Here, we test this assumption in the setting of gender inference. Using a controlled pronoun selection task, we introduce minimal, theoretically uninformative discourse context and find that this induces large, systematic shifts in model outputs. Correlations with cultural gender stereotypes, present in decontextualized settings, weaken or disappear once context is introduced, while theoretically irrelevant features, such as the gender of a pronoun for an unrelated referent, become the most informative predictors of model behavior. A Contextuality-by-Default analysis reveals that, in 19--52\% of cases across models, this dependence persists after accounting for all marginal effects of context on individual outputs and cannot be attributed to simple pronoun repetition. These findings show that LLM outputs violate contextual invariance even under near-identical syntactic formulations, with implications for bias benchmarking and deployment in high-stakes settings.

大模型偏见上下文依赖性别刻板印象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。