arXiv:2605.02930cs.NEcs.LG2026-05

用进化方法解析大模型关系,揭示关键训练数据与层的作用

Analysis and Explainability of LLMs Via Evolutionary Methods

论文配图:Analysis and Explainability of LLMs Via Evolutionary Methods
图 1 · 摘自论文原文
  • 将模型权重类比基因型、输出文本类比表型,构建进化分析框架
  • 实验中进化树准确还原真实训练树结构,识别出对模型影响最大的权重层
  • 无需标注即可构建黑箱模型的进化树,适合研究模型演化与数据贡献

进化方法在遗传学、生物学和生态学等领域长期用于分析与解释。本文将这些方法拓展至神经网络,特别是大语言模型(LLMs),以更好地分析和解释模型间的关系。通过将模型权重类比为基因型、输出文本类比为表型,我们提升了对模型谱系、重要数据集、不同模型层角色以及模型关系可视化的理解。在受控实验中,我们估计的进化树能可靠恢复真实训练树的拓扑结构。进一步地,根据权重差异识别出最重要的权重层,并通过表型实验发现一个训练数据集似乎比其他数据集提供更有效的信息。最后,我们生成了黑箱基础模型的无监督进化树。整个过程中,我们提供了可视化支持,使模型间进化关系的理解更加清晰。

原文摘要 · Abstract (English)

Evolutionary methods have long been useful for analysis and explanation in genetics, biology, ecology, and related fields. In this work, we extend these methods to neural networks, specifically large language models (LLMs), to better analyze and explain relationships among models. We show how relating weights to genotypes and output text to phenotypes can improve our understanding of model lineage, important datasets, the roles of different model layers, and visualization of model relationships. We demonstrate this in a controlled experiment, where our estimated evolutionary trees reliably recover the topology of the ground-truth training tree. We further identify the most important weight layers according to weight differences and show through phenotypic experiments that one training dataset appears to contribute more useful information than the others. Finally, we generate an unsupervised evolutionary tree of black-box foundation models. Throughout, we provide visualizations that support a clearer understanding of evolutionary relationships among LLMs.

大模型分析进化计算可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。