arXiv:2606.27632cs.CL2026-06

Yuvion LLM专为对抗性安全设计,提升模型在恶意攻击下的可靠性。

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

论文配图:Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety
图 1 · 摘自论文原文
  • 构建对抗感知数据,融合风险监督与强化学习优化安全策略
  • 在93项评测中表现优异,8B版本超越GPT-5.4等更大模型
  • 适合需要高安全性的实际应用,如内容审核与智能代理系统

随着大语言模型在真实系统中的广泛应用,安全失效仍可能导致有害输出和滥用。我们提出,安全的本质是对抗性的:许多问题并非源于自然输入,而是来自刻意规避模型政策的攻击行为。然而现有通用模型开发普遍忽视这一特性,在涉及规划、工具使用和多步推理的真实场景中,安全性能评估常高估实际鲁棒性。为此,我们提出Yuvion LLM,一个以对抗鲁棒性与智能体能力为核心目标的大语言模型。其流程包括对抗感知数据构建、知识增强的持续预训练,以及基于策略的安全多任务后训练,涵盖风险感知的监督微调、基于强化学习的策略优化,以及面向复杂安全场景的工具使用与多步推理的安全感知强化学习。我们还推出了Yuvion LLM RiskEval(YLRE),包含93个跨四类评估的基准,覆盖多样开放与内部测试,聚焦安全、对抗鲁棒性与现实能力需求。在这些评测中,Yuvion LLM在安全相关指标上表现突出,尤其在对抗条件下展现强鲁棒性,同时保持良好整体能力。值得注意的是,Yuvion-8B在多个安全任务上优于多数顶尖基线,包括显著更大的GPT-5.4和Qwen3-MAX。

原文摘要 · Abstract (English)

As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safety is adversarial: many failures arise not from natural inputs alone, but from strategic attempts to evade model policies and safeguards. However, existing general-purpose model development largely overlook this adversarial nature, and often remain insufficient for realistic safety scenarios involving planning, tool use, and multi-step reasoning, causing measured safety performance to overestimate real deployment robustness. To address this gap, we present Yuvion LLM, a large language model built for adversarially robust content safety and broader AI safety. Yuvion LLM treats adversarial robustness and agentic capability as first-class objectives. Its pipeline combines adversarially aware data construction, knowledge-enhanced continued pretraining, and policy-grounded multi-task safety post-training, including risk-aware supervised fine-tuning and reinforcement learning-based policy optimization, together with safety-aware agentic reinforcement learning for tool use and multi-step reasoning in complex safety scenarios. We further introduce the Yuvion LLM RiskEval (YLRE), a collection of 93 benchmarks across four evaluation categories, covering diverse open and internal evaluations with a focus on safety, adversarial robustness, and real-world capability requirements. Across these evaluations, Yuvion LLM demonstrates clear advantages on safety-focused benchmarks and particularly strong robustness under adversarial conditions, while maintaining solid overall capability. Notably, Yuvion-8B outperforms most state-of-the-art baselines, including substantially larger models such as GPT-5.4 and Qwen3-MAX, on several safety tasks.

对抗鲁棒内容安全智能体安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。