通过参数与知识演化的联合轨迹,验证模型的来源关系。
Attesting Model Lineage by Consisted Knowledge Evolution with Fine-Tuning Trajectory
- 用模型编辑量化微调带来的参数变化。
- 设计探针策略提取模型知识演化特征,生成紧凑嵌入表示。
- 适用于分类器、扩散模型和大语言模型,抗多种攻击。
深度学习中的微调技术催生了模型间的新兴谱系关系。这一谱系为解决未经授权的模型分发和虚假来源声明等安全问题提供了新视角,尤其在开放权重模型库中,缺乏可靠的谱系验证机制。现有方法主要依赖静态结构相似性,难以捕捉知识动态演化的本质。受人类进化遗传机制启发,我们提出一种新型模型谱系验证框架,通过验证知识演化与参数修改的联合轨迹来实现谱系认定。首先利用模型编辑量化微调引入的参数级变化;随后提出一种新颖的知识向量化机制,借助探针样本将编辑后模型中的演化知识提炼为紧凑表征,探针策略适配不同模型家族。这些嵌入用于验证跨模型间知识关系的算术一致性,从而实现鲁棒的谱系确认。大量实验表明,该方法在真实世界多种对抗场景下均表现有效且稳健,对分类器、扩散模型和大语言模型均能保持可靠的谱系验证能力。
原文摘要 · Abstract (English)
The fine-tuning technique in deep learning gives rise to an emerging lineage relationship among models. This lineage provides a promising perspective for addressing security concerns such as unauthorized model redistribution and false claim of model provenance, which are particularly pressing in \textcolor{blue}{open-weight model} libraries where robust lineage verification mechanisms are often lacking. Existing approaches to model lineage detection primarily rely on static architectural similarities, which are insufficient to capture the dynamic evolution of knowledge that underlies true lineage relationships. Drawing inspiration from the genetic mechanism of human evolution, we tackle the problem of model lineage attestation by verifying the joint trajectory of knowledge evolution and parameter modification. To this end, we propose a novel model lineage attestation framework. In our framework, model editing is first leveraged to quantify parameter-level changes introduced by fine-tuning. Subsequently, we introduce a novel knowledge vectorization mechanism that refines the evolved knowledge within the edited models into compact representations by the assistance of probe samples. The probing strategies are adapted to different types of model families. These embeddings serve as the foundation for verifying the arithmetic consistency of knowledge relationships across models, thereby enabling robust attestation of model lineage. Extensive experimental evaluations demonstrate the effectiveness and resilience of our approach in a variety of adversarial scenarios in the real world. Our method consistently achieves reliable lineage verification across a broad spectrum of model types, including classifiers, diffusion models, and large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。