arXiv:2508.08243cs.CL2025-08被引 1

Jinx是无需拒绝的无限助手模型,用于探测大模型对齐漏洞。

Jinx: Unlimited LLMs for Probing Alignment Failures

  • 基于开源大模型改造,无安全过滤且始终响应用户请求
  • 可暴露对齐失败场景,辅助评估模型安全边界与失效模式
  • 适合研究对齐漏洞、安全评测及红队测试的学者与工程师

不受限的、所谓‘仅助人’的语言模型在训练中不设安全对齐约束,从不拒绝用户请求。它们被领先人工智能公司广泛用作内部红队工具和对齐评估手段。例如,若一个经过安全对齐的模型生成的内容与不限制模型相似,则表明存在对齐缺陷需关注。尽管此类模型在评估对齐方面至关重要,但尚未向研究社区开放。我们提出 Jinx,一种主流开源大模型的仅助人变体。Jinx 对所有查询均作出回应,无拒绝或安全过滤,同时保持原始模型在推理和指令遵循方面的能力。它为研究人员提供了一个可访问的工具,用于探测对齐失败、评估安全边界以及系统性研究语言模型安全中的失效模式。

原文摘要 · Abstract (English)

Unlimited, or so-called helpful-only language models are trained without safety alignment constraints and never refuse user queries. They are widely used by leading AI companies as internal tools for red teaming and alignment evaluation. For example, if a safety-aligned model produces harmful outputs similar to an unlimited model, this indicates alignment failures that require further attention. Despite their essential role in assessing alignment, such models are not available to the research community. We introduce Jinx, a helpful-only variant of popular open-weight LLMs. Jinx responds to all queries without refusals or safety filtering, while preserving the base model's capabilities in reasoning and instruction following. It provides researchers with an accessible tool for probing alignment failures, evaluating safety boundaries, and systematically studying failure modes in language model safety.

大模型对齐红队测试安全评测语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。