Yuvion VL专攻多模态安全,能识别对抗性内容并提升模型鲁棒性。
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

- 构建对抗性数据流水线,融合领域知识与推理标注,生成高质量多模态样本。
- 三阶段训练提升安全任务表现,32B模型超越主流开源与闭源模型。
- 引入混淆-对比微调,精准区分视觉相似但安全含义不同的案例。
通用模型在识别真实世界多模态风险时常表现不稳定,主要因内容与AI安全具有固有的多模态对抗特性。我们提出Yuvion VL,一个专为内容与AI安全设计的多模态大语言模型家族,包含指令微调和推理导向两种变体。该模型将安全视为本质对抗性问题,从数据构建到训练全程强化对抗鲁棒性。数据方面,开发自动化流水线,结合对抗感知数据合成与多阶段质量控制,生成大规模、高质的多模态样本,附带领域知识与推理标注。训练采用三阶段流程:持续预训练以对齐风险概念的跨模态理解,指令后训练适配生产级安全任务,推理后训练增强复杂任务中的可解释性与性能。进一步提出“混淆-对比微调”(Confuse-then-Contrast Fine-Tuning),通过挖掘模型特定混淆点,构建多图像对比组,强制区分细微的视觉-语义差异,在对抗性安全任务中实现精准判别。为支持严谨评估,引入Yuvion VL RiskEval(YVRE)基准集,涵盖开放与内部评估,聚焦内容安全、对抗鲁棒性与真实场景能力。实验表明,Yuvion VL-32B在安全性能上达到行业领先水平,优于同等规模的开源模型及顶级闭源商业模型,同时保持相当的通用能力。
原文摘要 · Abstract (English)
General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI safety. We present Yuvion VL, a family of multimodal large language models purpose-built for content and AI safety, with both instruction-tuned and reasoning-oriented variants. Yuvion VL addresses this gap by treating safety as an inherently adversarial and multimodal problem and designing the entire pipeline around adversarial robustness. For data construction, we develop an automated pipeline integrating adversarial-aware data synthesis with multi-stage quality control, producing large-scale, high-quality multimodal samples augmented with domain knowledge and reasoning annotations. For training, we adopt a three-stage pipeline that includes continued pretraining for risk-concept cross-modal alignment, instruct post-training for production-grade safety tasks, and reasoning post-training for enhanced interpretability and performance in complex tasks. We further introduce Confuse-then-Contrast Fine-Tuning, a contrastive framework that mines model-specific confusions and constructs multi-image contrastive groups to enforce explicit discrimination of fine-grained visual-semantic elements, enabling the model to distinguish between visually similar cases with different safety implications in adversarial safety tasks. To support rigorous evaluation, we further introduce Yuvion VL RiskEval (YVRE), a collection of benchmarks covering diverse open and internal evaluations, with a focus on content and AI safety, adversarial robustness, and real-world capability requirements. Experiments show that Yuvion VL-32B achieves industry-leading safety performance, surpassing comparably sized open-source models and best closed-source commercial models, while maintaining comparable general capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。