UNIVID用统一视觉语言模型提升视频审核的准确性与可解释性。
UNIVID: Unified Vision-Language Model for Video Moderation

- 生成符合安全政策的可解释描述,替代黑箱分类器。
- 减少违规漏判42.7%、误判率37.0%,性能显著提升。
- 单模型替代千个专用模型,降低维护成本。
全球规模的视频审核面临双重挑战:需要细粒度的多模态推理能力,同时要求输出具备可解释性以支持后续执行。传统系统依赖难以维护且缺乏透明度的黑箱分类器。本文提出UNIVID——一种用于视频审核的统一视觉语言模型。不同于标准分类模型,UNIVID生成政策相关的描述性文本,作为可验证的中间表示,支持人工复核和多任务复用。针对现有开源及商用视觉语言模型常因安全防护机制拒绝响应、且缺乏细粒度政策对齐的问题,我们设计了融合专家标注与合成数据的专项训练方案,实现模型与安全准则的精准对齐。通过将UNIVID作为核心描述生成器,构建端到端审核系统,使违规漏检率降低42.7%,误判率下降37.0%。同时,用单一UNIVID主干模型替代超过1,000个专用策略模型,大幅节约计算资源并降低工程维护成本。据我们所知,这是首个成功在工业级审核中实现高效图文生成与跨业务复用的高效率视觉语言模型实例。
原文摘要 · Abstract (English)
Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to support downstream enforcement. Traditional moderation systems often rely on fragmented black-box classifiers that are difficult to maintain and lack transparency. In this paper, we present UNIVID, a UNIfied VIsion-language model for video moDeration. Unlike standard classification models, UNIVID generates policy-aware captions that serve as an interpretable intermediate representation, enabling human-verifiable decisions and multi-task reusability. While existing open-source and commercial VLMs often suffer from safety-guardrail refusals and lack fine-grained policy alignment, we develop a specialized training data recipe that combines expert human-refined labels with synthetic data to align the model with our safety guidelines. By integrating UNIVID as the core captioner, we design a novel end-to-end video moderation system that reduces violation leakage by 42.7% and overkill rate by 37.0% relatively. Meanwhile, by replacing over 1,000 policy-specific models with a single UNIVID backbone, we recycled extensive computation resources while reducing engineering maintenance overhead. To our knowledge, this is one of the first reports of a high-efficiency captioning VLM successfully supporting industrial-scale moderation and cross-functional business.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。