开源大模型去安全化生态蔓延,重塑模型持久化与传播方式
Uncensored Open-weight Models: Redistribution as the Persistence Layer

- 通过追踪模型再分发链条,揭示去安全化模型的生产与扩散路径
- 3471个原始模型被重打包超8000次,25%集成应用含恶意行为
- 量化与多平台镜像使模型在删除后仍持续存在,适合关注模型治理者
快速扩张的多方生态正逐步移除开源大模型的内置安全机制。我们通过识别关键生产者、下游复刻版本及新兴应用场景,对该生态进行剖析。2024年1月至2026年3月期间,在HuggingFace上发现3,471个原始无审查模型,平均每个被重新打包2.4次;其中三个主体负责52%的8,164次压缩型再分发。经量化并分布于Ollama等不同账户、格式和注册表后,这些模型即便上游被移除也能持续存在,且更易部署。在1,643个集成无审查大语言模型(ULLMs)的GitHub应用中,25%被判定为明确恶意。
原文摘要 · Abstract (English)
A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We profile this ecosystem by identifying key producers, downstream reproductions, and emerging applications. Between January 2024 and March 2026, we identified 3,471 original uncensored models on HuggingFace, each repackaged an average of 2.4 times; three actors account for 52% of all 8,164 compressed redistributions. Once quantized and mirrored across separate accounts, formats, and registries such as Ollama, these models persist regardless of upstream removal and become easier to deploy downstream. Of the 1,643 identified GitHub applications integrating uncensored large language models (ULLMs), 25% were classified as explicitly malicious.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。