模型共识程度取决于评估方式,没有绝对标准。
The Subjectivity of Monoculture
- 用不同基准模型判断模型是否过度一致,结果差异显著。
- 同一模型在不同问题集上表现可能从高度相关变为独立。
- 研究揭示了模型共识评估的主观性,适合方法论反思者。
机器学习模型(包括大语言模型)常被认为存在单一化现象,即输出高度一致。但模型一致性过高究竟意味着什么?我们指出这一问题本质上是主观的,依赖两个关键决策:首先,分析者需设定一个用于衡量“独立性”的基准模型,而该选择本身具有主观性,且不同的基准会导致对过度一致性的截然不同推断;其次,推断结果还取决于所考察的模型群体与问题集合。在不同问题集或不同同行模型下,同一模型可能表现出高度相关或完全独立。在两个大规模基准上的实验验证了理论发现,例如,使用包含题目难度的基准模型时,推断结果与以往不考虑难度的工作有显著差异。综上,我们的研究将单一化评估重新定义为一种依赖上下文的推断问题,而非模型行为的绝对属性。
原文摘要 · Abstract (English)
Machine learning models -- including large language models (LLMs) -- are often said to exhibit monoculture, where outputs agree strikingly often. But what does it actually mean for models to agree too much? We argue that this question is inherently subjective, relying on two key decisions. First, the analyst must specify a baseline null model for what "independence" should look like. This choice is inherently subjective, and as we show, different null models result in dramatically different inferences about excess agreement. Second, we show that inferences depend on the population of models and items under consideration. Models that seem highly correlated in one context may appear independent when evaluated on a different set of questions, or against a different set of peers. Experiments on two large-scale benchmarks validate our theoretical findings. For example, we find drastically different inferences when using a null model with item difficulty compared to previous works that do not. Together, our results reframe monoculture evaluation not as an absolute property of model behavior, but as a context-dependent inference problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。