arXiv:2504.04216cs.CL2025-04

用困惑度与曲率差评估大模型相似性,防抄袭更精准

A Perplexity and Menger Curvature-Based Approach for Similarity Evaluation of Large Language Models

  • 结合困惑度曲线和赵尔曲率差异量化模型相似性
  • 在多类模型与领域上均优于传统方法
  • 适合检测模型复制,保护大模型知识产权

大语言模型(LLMs)的兴起引发了关于版权侵权和数据及模型使用不当的担忧。例如,对现有模型进行微小修改可能被用来虚假宣称开发了新模型,导致模型抄袭和所有权侵犯问题。本文提出一种新度量方法,通过困惑度曲线和赵尔曲率差异来量化大模型的相似性。全面实验验证了该方法的有效性,证明其在多个模型和领域中均优于基线方法,并具备良好的泛化能力。此外,通过模拟实验展示了该方法在检测模型复制方面的潜力,强调其在维护大模型原创性和完整性方面的价值。代码已开源:https://github.com/zyttt-coder/LLM_similarity。

原文摘要 · Abstract (English)

The rise of Large Language Models (LLMs) has brought about concerns regarding copyright infringement and unethical practices in data and model usage. For instance, slight modifications to existing LLMs may be used to falsely claim the development of new models, leading to issues of model copying and violations of ownership rights. This paper addresses these challenges by introducing a novel metric for quantifying LLM similarity, which leverages perplexity curves and differences in Menger curvature. Comprehensive experiments validate the performance of our methodology, demonstrating its superiority over baseline methods and its ability to generalize across diverse models and domains. Furthermore, we highlight the capability of our approach in detecting model replication through simulations, emphasizing its potential to preserve the originality and integrity of LLMs. Code is available at https://github.com/zyttt-coder/LLM_similarity.

大模型评估相似性检测版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。