用基因图谱追踪大模型演化,揭示隐藏的传承关系。
LLM DNA: Tracing Model Evolution via Functional Representations
- 将模型功能行为抽象为低维基因表示,数学上证明其可遗传性。
- 在305个模型上验证,能准确还原已知演化路径并发现新关系。
- 无需训练、适配任意架构,适合模型管理与溯源分析。
大规模语言模型(LLMs)数量激增,但其通过微调、蒸馏或适应产生的演化关系常未记录或难以辨识,给模型管理带来挑战。现有方法受限于任务特异性、固定模型集或对分词器、架构的严格假设。受生物DNA启发,我们数学定义了LLM DNA为功能行为的低维、双李普希茨表示,证明其具备继承性与遗传决定性,并确立其存在性。基于此理论,构建了一套通用、可扩展、无需训练的基因提取流程。在305个模型上的实验表明,该方法与小规模已有研究结果一致,且在特定任务上表现更优或具竞争力。此外,基因比对揭示了此前未知的模型关联。进一步利用系统发育算法构建了模型演化树,其结果与从编码器-解码器向解码器仅架构的转变趋势相符,反映时间演进规律,并揭示不同模型家族演化速度差异。
原文摘要 · Abstract (English)
The explosive growth of large language models (LLMs) has created a vast but opaque landscape: millions of models exist, yet their evolutionary relationships through fine-tuning, distillation, or adaptation are often undocumented or unclear, complicating LLM management. Existing methods are limited by task specificity, fixed model sets, or strict assumptions about tokenizers or architectures. Inspired by biological DNA, we address these limitations by mathematically defining LLM DNA as a low-dimensional, bi-Lipschitz representation of functional behavior. We prove that LLM DNA satisfies inheritance and genetic determinism properties and establish the existence of DNA. Building on this theory, we derive a general, scalable, training-free pipeline for DNA extraction. In experiments across 305 LLMs, DNA aligns with prior studies on limited subsets and achieves superior or competitive performance on specific tasks. Beyond these tasks, DNA comparisons uncover previously undocumented relationships among LLMs. We further construct the evolutionary tree of LLMs using phylogenetic algorithms, which align with shifts from encoder-decoder to decoder-only architectures, reflect temporal progression, and reveal distinct evolutionary speeds across LLM families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。