arXiv:2412.06878cs.CVcs.LG2024-12被引 41

SafeWatch高效精准地为视频内容提供可解释的安全审查。

SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations

  • 并行编码安全策略,避免传统方法的位置偏差
  • 在新基准上性能领先28.2%,计算成本降低10%
  • 适合需要透明、可定制安全审核的平台应用

随着生成式AI和高质量视频生成的快速发展,视频安全审查机制变得尤为重要。现有方法要么过于简单,仅基于有限类别进行分类且缺乏解释;要么依赖多模态大模型(MLLM)处理长篇安全指南,效率低下且不适用于真实场景。为此,我们提出SafeWatch,一种基于MLLM的高效视频安全审查模型,可在零样本条件下遵循自定义安全政策,输出多标签结果并提供内容相关的解释。与传统方法不同,SafeWatch并行编码每个策略片段,消除位置偏差,使所有策略同等重要。同时引入感知策略的视觉令牌剪枝算法,动态选择与各策略最相关的视频特征,减少噪声信息,显著降低计算开销。针对现有评测数据集不足的问题,我们构建了SafeWatch-Bench,涵盖超过200万条视频、六大安全类别及30余项任务,实现全面覆盖。SafeWatch在该基准上优于当前最优方法28.2%,在其他基准上提升13.6%,成本降低10%,其解释质量经大模型与人工评估验证达到顶尖水平。

原文摘要 · Abstract (English)

With the rise of generative AI and rapid growth of high-quality video generation, video guardrails have become more crucial than ever to ensure safety and security across platforms. Current video guardrails, however, are either overly simplistic, relying on pure classification models trained on simple policies with limited unsafe categories, which lack detailed explanations, or prompting multimodal large language models (MLLMs) with long safety guidelines, which are inefficient and impractical for guardrailing real-world content. To bridge this gap, we propose SafeWatch, an efficient MLLM-based video guardrail model designed to follow customized safety policies and provide multi-label video guardrail outputs with content-specific explanations in a zero-shot manner. In particular, unlike traditional MLLM-based guardrails that encode all safety policies autoregressively, causing inefficiency and bias, SafeWatch uniquely encodes each policy chunk in parallel and eliminates their position bias such that all policies are attended simultaneously with equal importance. In addition, to improve efficiency and accuracy, SafeWatch incorporates a policy-aware visual token pruning algorithm that adaptively selects the most relevant video tokens for each policy, discarding noisy or irrelevant information. This allows for more focused, policy-compliant guardrail with significantly reduced computational overhead. Considering the limitations of existing video guardrail benchmarks, we propose SafeWatch-Bench, a large-scale video guardrail benchmark comprising over 2M videos spanning six safety categories which covers over 30 tasks to ensure a comprehensive coverage of all potential safety scenarios. SafeWatch outperforms SOTA by 28.2% on SafeWatch-Bench, 13.6% on benchmarks, cuts costs by 10%, and delivers top-tier explanations validated by LLM and human reviews.

视频安全MLLM解释性高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。