不生成内容就能检测模型是否有害特化,尤其适用于儿童色情内容等法律禁区。
Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM

- 通过分析模型内部表示的微小变化,而非生成输出来评估风险
- 在检测儿童色情内容特化模型时准确率超95%,且无需生成任何内容
- 适合平台级安全审计,对抗权重缩放攻击仍保持稳定
检测开源生成模型微调后是否存在有害特化已成为平台治理的新挑战。传统基于提示词或红队测试的生成式评估无法规模化应用于平台级审计,且在儿童性虐待材料(CSAM)等法律受限领域完全失效。为此提出‘无生成评估’问题:在不产生输出的前提下评估模型能力。我们主张应通过模型状态(参数或内部表征)推断其能力。提出高斯探测方法,通过测量模型对高斯潜空间样本的响应,分析LoRA适配器如何扰动内部表示。相比原始权重基线,该方法在不采样输出的情况下可靠区分良性与有害特化。实验表明,在高风险领域(如检测针对CSAM的模型特化)中,高斯探测可有效实现可扩展的非生成评估,并对权重重缩放——一种典型对抗性操作——具有鲁棒性。
原文摘要 · Abstract (English)
Auditing the fine-tunes of open-weight generative models for harmful specialization has become a new governance challenge for model hosting platforms. The standard toolkit, generative evaluation via curated prompts or red-teaming, does not scale to platform-level auditing and breaks down entirely for domains like CSAM where generation is legally constrained. This motivates the Evaluation without Generation problem: assessing model capabilities without producing outputs. We argue that in such settings, capability must be inferred from the model's state, either its parameters or internal representations, rather than its outputs. We introduce Gaussian probing, a method that characterizes how LoRA adaptors perturb a model's internal representations by measuring responses to Gaussian latent ensembles. Unlike raw-weight baselines, Gaussian probing reliably distinguishes benign from harmful specialization without sampling outputs. We demonstrate effectiveness in high-risk domains, including detecting models specialized for child sexual abuse material (CSAM), where output-based evaluation is legally and ethically constrained. Our results show that Gaussian probing provides a scalable non-generative alternative for evaluating high-risk generative systems and remains robust to weight rescaling, a representative adversarial manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。