arXiv:2606.21645cs.CLcs.LG2026-06

研究大模型对词语搭配顺序的偏好,发现其行为符合习惯但统计细节不精准。

Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models

论文配图:Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models
图 1 · 摘自论文原文
  • 将词语搭配顺序视为分布对齐问题,构建跨8语言600对数据集
  • 模型虽能复现主流顺序,但整体分布匹配度不高
  • 可透过中间层探针测量并操控模型的搭配偏好

大型语言模型(LLMs)能自然重现常见表达,但其对频率分布的建模能力仍不明确。本文以‘男女’等语言双词搭配为例,这些组合语法均合法,但跨语言呈现不同习惯性顺序。我们将双词顺序建模为分布对齐问题,构建了涵盖8种语言、共600对双词搭配的多语言数据集。通过分类与分布度量方法,比较6个开源大模型生成的顺序概率与语料库实际偏好。结果显示,模型在强惯例化搭配上行为上能复现主流顺序,但整体分布对齐程度不足,表明模型看似确定的顺序偏好,实际上未能忠实捕捉语言使用的统计细微差别。稀疏探针验证显示,偏好强度概念部分编码于中到深层网络,沿探针方向调节可改变模型生成的顺序分布,证明大模型的统计偏好可通过内部表示被机制性测量与干预。

原文摘要 · Abstract (English)

Large language models (LLMs) can readily reproduce conventional expressions, yet their ability to model gradient frequency distributions remains underexplored. We investigate this using linguistic binomials, such as men and women, where both word permutations are grammatically valid but exhibit distinct, cross-linguistic variations in conventionality. We formalize binomial ordering as a distributional alignment problem, and construct a multilingual dataset of 600 binomial pairs across 8 languages. With categorical and distributional metrics, we measure and compare the corpus-derived preferences with model-induced ordering probabilities of 6 open-weight LLMs. While models often behaviorally recover the dominant corpus-preferred order, particularly for strongly conventionalized pairs, they align less well with the exact corpus preference distributions. This suggests that apparent directional order overstates how faithfully LLMs capture the statistical nuances of language use. Sparse probing verifies that the concept of preference strength is partially encoded among middle-to-late layers, and steering along probe-derived directions alters model-induced ordering distributions, demonstrating that the statistical behavioral preference of LLMs can be mechanistically measured and manipulated via internal representations.

语言模型搭配顺序分布建模探针分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。