用文本描述桥接视频安全对齐,提升视觉语言模型安全性
SafeVid: Toward Safety Aligned Video Large Multimodal Models
- 以视频文本描述为桥梁,实现文本安全规则向视频域迁移
- 构建35万对视频安全偏好数据集,使模型安全得分提升42.39%
- 适合关注视频生成安全、多模态模型对齐的研究者使用
随着视频大型多模态模型(VLMMs)快速发展,其内在复杂性带来显著安全挑战,尤其体现在静态安全对齐无法适配动态视频场景的泛化不匹配问题。本文提出SafeVid框架,通过详细视频文本描述作为解释性桥梁,将稳健的文本安全对齐能力迁移至视频领域,实现基于大模型规则驱动的安全推理。该框架包含三个环节:1)构建首个35万对的视频特定安全偏好数据集SafeVid-350K;2)采用直接偏好优化(DPO)对VLMM进行针对性对齐;3)通过新提出的SafeVidBench基准进行全面评估。实验表明,与SafeVid-350K对齐后,如LLaVA-NeXT-Video等模型在SafeVidBench上的安全表现显著提升,最高达42.39%。SafeVid提供关键资源与系统化方法,证明利用文本描述作为安全推理通道能显著增强VLMM的安全对齐效果。相关数据集已公开于HuggingFace。
原文摘要 · Abstract (English)
As Video Large Multimodal Models (VLMMs) rapidly advance, their inherent complexity introduces significant safety challenges, particularly the issue of mismatched generalization where static safety alignments fail to transfer to dynamic video contexts. We introduce SafeVid, a framework designed to instill video-specific safety principles in VLMMs. SafeVid uniquely transfers robust textual safety alignment capabilities to the video domain by employing detailed textual video descriptions as an interpretive bridge, facilitating LLM-based rule-driven safety reasoning. This is achieved through a closed-loop system comprising: 1) generation of SafeVid-350K, a novel 350,000-pair video-specific safety preference dataset; 2) targeted alignment of VLMMs using Direct Preference Optimization (DPO); and 3) comprehensive evaluation via our new SafeVidBench benchmark. Alignment with SafeVid-350K significantly enhances VLMM safety, with models like LLaVA-NeXT-Video demonstrating substantial improvements (e.g., up to 42.39%) on SafeVidBench. SafeVid provides critical resources and a structured approach, demonstrating that leveraging textual descriptions as a conduit for safety reasoning markedly improves the safety alignment of VLMMs. We have made SafeVid-350K dataset (https://huggingface.co/datasets/yxwang/SafeVid-350K) publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。