arXiv:2503.04332cs.CRcs.LG2025-03被引 12

提出新方法可追踪被微调的黑箱大模型来源,破解非法使用难题。

The Challenge of Identifying the Origin of Black-Box Large Language Models

  • 通过优化连续空间中的对抗性标记嵌入,主动植入追踪信号。
  • 在30个模型和2个真实API上验证,对微调模型识别率显著提升。
  • 适合关注模型版权保护与合规审计的研究者与政策制定者。

大语言模型巨大的商业潜力引发了对其未经授权使用的担忧。第三方可通过微调定制模型并仅提供黑箱API访问,从而隐藏非法使用行为,使外部审计变得困难。这不仅加剧了不公平竞争,还违反了许可协议。为此,识别黑箱大模型的来源成为根本解决方案。本文通过实验揭示了现有被动与主动识别方法在30个大模型及两个真实黑箱API上的局限性。随后提出主动识别技术PlugAE,通过在连续空间中优化对抗性标记嵌入,并主动注入模型以实现溯源与识别。实验表明,PlugAE在识别微调衍生模型方面有显著提升。此外,论文呼吁建立法律框架以更好应对大模型未经授权使用的挑战。

原文摘要 · Abstract (English)

The tremendous commercial potential of large language models (LLMs) has heightened concerns about their unauthorized use. Third parties can customize LLMs through fine-tuning and offer only black-box API access, effectively concealing unauthorized usage and complicating external auditing processes. This practice not only exacerbates unfair competition, but also violates licensing agreements. In response, identifying the origin of black-box LLMs is an intrinsic solution to this issue. In this paper, we first reveal the limitations of state-of-the-art passive and proactive identification methods with experiments on 30 LLMs and two real-world black-box APIs. Then, we propose the proactive technique, PlugAE, which optimizes adversarial token embeddings in a continuous space and proactively plugs them into the LLM for tracing and identification. The experiments show that PlugAE can achieve substantial improvement in identifying fine-tuned derivatives. We further advocate for legal frameworks and regulations to better address the challenges posed by the unauthorized use of LLMs.

模型溯源黑箱检测版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。