Llama Guard 3 Vision可检测图文对话中的有害内容,提升AI安全性。
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
- 基于Llama 3.2-Vision微调,支持图文输入与输出的有害内容识别。
- 在MLCommons测试中表现优异,能有效识别多模态有害提示与回复。
- 适用于需要图文安全过滤的AI对话系统,适合开发者部署使用。
我们提出Llama Guard 3 Vision,一种基于多模态大模型的人机对话安全保障工具,支持图像理解场景下的内容防护,可用于多模态大模型的输入(提示分类)与输出(响应分类)。与此前仅处理文本的Llama Guard版本不同,该模型专为图像推理任务设计,优化于检测包含文本与图像的有害提示及对这些提示的文本回复。模型在Llama 3.2-Vision基础上微调,在内部基于MLCommons分类体系的基准测试中表现良好,并验证了其对对抗攻击的鲁棒性。我们认为,Llama Guard 3 Vision是构建更强大、更可靠的多模态人机对话内容审核工具的良好起点。
原文摘要 · Abstract (English)
We introduce Llama Guard 3 Vision, a multimodal LLM-based safeguard for human-AI conversations that involves image understanding: it can be used to safeguard content for both multimodal LLM inputs (prompt classification) and outputs (response classification). Unlike the previous text-only Llama Guard versions (Inan et al., 2023; Llama Team, 2024b,a), it is specifically designed to support image reasoning use cases and is optimized to detect harmful multimodal (text and image) prompts and text responses to these prompts. Llama Guard 3 Vision is fine-tuned on Llama 3.2-Vision and demonstrates strong performance on the internal benchmarks using the MLCommons taxonomy. We also test its robustness against adversarial attacks. We believe that Llama Guard 3 Vision serves as a good starting point to build more capable and robust content moderation tools for human-AI conversation with multimodal capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。