模型权重可压缩性成安全漏洞,压缩后传输时间从月缩至天
Aggressive Compression Enables LLM Weight Theft
- 针对性压缩放松解压要求,实现16到100倍压缩率
- 压缩后传输时间从数月缩短至数日,显著提升窃取效率
- 水印追踪是低成本高效防御方案,适合实际部署
随着前沿大模型愈发强大且研发成本高昂,攻击者窃取模型权重的动机增强。本文研究通过网络从数据中心窃取大语言模型权重的攻击方式。尽管此类攻击为多步骤行为,我们发现模型权重的可压缩性是决定泄露风险的关键因素。通过专为窃取设计的压缩策略,放宽解压约束,攻击者可实现16至100倍压缩,仅付出微小性能损失,使模型权重非法传输时间由数月缩短至数日。我们评估了三类防御策略:降低模型可压缩性、隐藏模型位置、使用数字水印追踪溯源。三者均有潜力,但水印方案兼具高效与低成本,是最具吸引力的防护手段。
原文摘要 · Abstract (English)
As frontier AIs become more powerful and costly to develop, adversaries have increasing incentives to steal model weights by mounting exfiltration attacks. In this work, we consider exfiltration attacks where an adversary attempts to sneak model weights out of a datacenter over a network. While exfiltration attacks are multi-step cyber attacks, we demonstrate that a single factor, the compressibility of model weights, significantly heightens exfiltration risk for large language models (LLMs). We tailor compression specifically for exfiltration by relaxing decompression constraints and demonstrate that attackers could achieve 16x to 100x compression with minimal trade-offs, reducing the time it would take for an attacker to illicitly transmit model weights from the defender's server from months to days. Finally, we study defenses designed to reduce exfiltration risk in three distinct ways: making models harder to compress, making them harder to 'find,' and tracking provenance for post-attack analysis using forensic watermarks. While all defenses are promising, the forensic watermark defense is both effective and cheap, and therefore is a particularly attractive lever for mitigating weight-exfiltration risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。