arXiv:2606.16407cs.CLcs.LG2026-06

揭示大模型指代一致性背后的三种内在机制及其竞争关系。

A Mechanistic Understanding of Pronoun Fidelity in LLMs

论文配图:A Mechanistic Understanding of Pronoun Fidelity in LLMs
图 1 · 摘自论文原文
  • 从模型内部出发,识别出实体绑定、最近性偏倚和刻板印象三类因果机制。
  • 三者共同解释91%至99.5%的指代行为,无单一机制能完全说明。
  • 发现两种复制路径:概念级与词元级,反映不同机制的竞争机制。

指代的一致性和鲁棒性对生成内容的公平性与连贯性至关重要,但当多个指代对象使用不同代词时,大语言模型表现不佳。以往研究仅依赖行为分析,无法反映模型内部运作。本文从机制层面切入,检验三个机制——实体绑定(G)、最近性偏倚(R)和刻板印象偏倚(S)——是否在多个前沿语言模型中被因果实现。通过边界分布式对齐搜索,发现三者作为分布在网络深度中的因果子空间共存。单一机制无法完全解释模型行为,但三者组合可一致解释91%至99.5%的现象。注意力头分析进一步揭示两条竞争性复制路径:实体绑定与刻板印象共享一个局部的概念级路径,用于检索绑定的职业-代词单元;而最近性偏倚则使用分布式的词元级路径,重复表面形式。综上,指代一致性源于同时活跃的因果子空间之间的竞争。

原文摘要 · Abstract (English)

Faithful and robust pronoun use is important for fair and coherent generations, yet large language models largely fail when multiple referents use different pronouns. To study the interplay of reasoning, repetition, and bias in this task, prior work relies exclusively on behavioural approaches, which may not reflect a model's internal workings. Therefore, we provide a mechanistic, model-internal perspective on pronoun fidelity, testing whether three mechanisms -- group entity binding (G), recency bias (R), and stereotypical bias (S) -- are causally implemented across several SOTA language models. Using Boundless Distributed Alignment Search, we find all three coexist as causal subspaces distributed across network depth. No single mechanism fully explains model behaviour, but a combination of the three consistently accounts for 91-99.5%. An attention head analysis further reveals two competing copying routes; group binding and stereotype share a localized concept-level route that retrieves a bound occupation-pronoun unit, while recency uses a distributed token-level route that repeats surface forms. In sum, pronoun fidelity arises from competition between simultaneously active causal subspaces.

指代一致性机制分析语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。