用生成过程差异预测大模型集体失效,比语义相似性更准。
Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

- 基于压缩距离衡量模型生成过程差异,避开语义表层相似。
- 38个模型中,过程多样性越高,模型对齐失效越少。
- 适合关注多模型系统安全性的研究者使用。
多样性是集体系统稳健性的关键因素,但其类型取决于系统特性和失效模式。对于多个语言模型组成的系统,即使行为和失败高度相关,仍常被视为独立组件。现有基于语义相似度的评估方法仅捕捉输出含义差异,未能反映深层生成过程差异。本文提出以归一化压缩距离(NCD)测量模型输出间的生成过程多样性,并通过置换控制残差化处理,提升可靠性。在38个语言模型上,该方法揭示了语义相似度无法捕捉的群体结构,且能有效预测跨10个独立基准任务中模型对的关联失效变异,其部分秩相关系数为-0.216(95%置信区间[-0.309, -0.122]),所有基准均呈负相关。结果表明,生成过程多样性越高,非语义相似性或能力差异导致的关联失效越低。该方法为安全相关的多模型系统多样性分析提供了新路径。
原文摘要 · Abstract (English)
Diversity is a widely observed factor in the resilient function of collective systems, yet the type of diversity that matters depends on the properties and failure modes of the system. This distinction is important for systems composed of multiple language models. Different models may be treated as independent components even when their behaviour and failures remain strongly correlated. Assessments of language-model populations using semantic similarity demonstrate limited semantic diversity, but this captures only differences in the meaning of observed outputs. We argue that a more fundamental notion of model diversity is generative-process diversity, the differences between processes capable of generating the observed outputs. Drawing from Algorithmic Information Theory, we use Normalised Compression Distance between raw model outputs, residualised against a permutation control, as a measure of inferred generative-process diversity. Across 38 language models, this measure identifies population structure missed by semantic similarity and predicts cross-task variation in chance-corrected correlated failure among model pairs across ten disjoint benchmark families, beyond semantic similarity and model-pair capability. The cross-benchmark partial rank association is $-0.216$ with a 95% interval of $[-0.309,-0.122]$, and the estimate is negative on all ten benchmarks. These results indicate that increased generative-process diversity is associated with reduced correlated failure in model pairs that is not attributable to semantic similarity or capability. Inferred generative-process diversity offers a novel and practical approach for investigating diversity of multi-model systems in safety-relevant contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。