arXiv:2509.21486cs.CV2025-09EMNLP被引 17

用推理增强的多模态模型统一识别短视频违规内容

Reasoning-Enhanced Domain-Adaptive Pretraining of Multimodal Large Language Models for Short Video Content Governance

  • 引入三阶段预训练提升模型对视频细节、问题定义和逻辑推理的理解
  • 零样本和微调下均显著提升检测性能,对新出现问题有强泛化能力
  • 适合需要高效、通用内容审核的平台或研究者参考

短视频平台发展迅速,不当内容识别日益重要。现有方法通常为每类问题单独训练小型分类模型,依赖大量人工标注数据且缺乏跨问题泛化能力。本文提出一种推理增强的多模态大语言模型(MLLM)统一预训练范式,用于统一检测不当内容。针对短视频内容与MLLM原始预训练数据间的分布差异及复杂问题定义,设计三个针对性预训练任务:(1) 标题生成(Caption),增强模型对视频细节的感知;(2) 视觉问答(VQA),深化模型对问题定义与标注规范的理解;(3) 思维链(CoT),提升模型推理能力。实验表明,该预训练方法在零样本和监督微调(SFT)设置下均显著提升性能,且对新出现的、未见过的问题展现出强大泛化能力。

原文摘要 · Abstract (English)

Short video platforms are evolving rapidly, making the identification of inappropriate content increasingly critical. Existing approaches typically train separate and small classification models for each type of issue, which requires extensive human-labeled data and lacks cross-issue generalization. We propose a reasoning-enhanced multimodal large language model (MLLM) pretraining paradigm for unified inappropriate content detection. To address the distribution gap between short video content and the original pretraining data of MLLMs, as well as the complex issue definitions, we introduce three targeted pretraining tasks: (1) \textit{Caption}, to enhance the MLLM's perception of video details; (2) \textit{Visual Question Answering (VQA)}, to deepen the MLLM's understanding of issue definitions and annotation guidelines; (3) \textit{Chain-of-Thought (CoT)}, to enhance the MLLM's reasoning capability. Experimental results show that our pretraining approach significantly improves the MLLM's performance in both zero-shot and supervised fine-tuning (SFT) settings. In addition, our pretrained model demonstrates strong generalization capabilities to emergent, previously unseen issues.

内容审核多模态推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。