arXiv:2411.05854cs.MMcs.AI2024-11被引 13

构建视频平台有害内容分类体系,用大模型替代人工标注。

Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators

  • 提出六类网络危害分类体系,覆盖信息、仇恨、成瘾等
  • GPT-4-Turbo在19,422条视频上优于众包标注员
  • 为平台内容安全提供可落地的多模态检测方案

短视频平台如YouTube、Instagram和TikTok全球用户达数十亿,但用户常接触各类有害内容,包括误导性信息、仇恨言论、成瘾诱导、点击诱饵、色情内容及身体伤害等。现有检测面临定义不一、人工标注资源有限且心理负担重等问题。本研究构建了涵盖信息、仇恨与骚扰、成瘾、点击诱饵、性及身体伤害共六类的在线危害分类体系,并验证多模态大语言模型(MLLM)作为可靠标注者的能力。基于19,422条YouTube视频,分析14帧图像、缩略图与文本元数据,以领域专家标注为金标准,对比众包标注员(Mturk)与GPT-4-Turbo的表现。结果显示,GPT-4-Turbo在二分类(有害/无害)与多标签分类任务中均优于人类标注者。方法上拓展了大模型在多模态、多标签场景的应用边界;实践上为平台有害内容识别与治理提供了标准化框架。

原文摘要 · Abstract (English)

Short video platforms, such as YouTube, Instagram, or TikTok, are used by billions of users globally. These platforms expose users to harmful content, ranging from clickbait or physical harms to misinformation or online hate. Yet, detecting harmful videos remains challenging due to an inconsistent understanding of what constitutes harm and limited resources and mental tolls involved in human annotation. As such, this study advances measures and methods to detect harm in video content. First, we develop a comprehensive taxonomy for online harm on video platforms, categorizing it into six categories: Information, Hate and harassment, Addictive, Clickbait, Sexual, and Physical harms. Next, we establish multimodal large language models as reliable annotators of harmful videos. We analyze 19,422 YouTube videos using 14 image frames, 1 thumbnail, and text metadata, comparing the accuracy of crowdworkers (Mturk) and GPT-4-Turbo with domain expert annotations serving as the gold standard. Our results demonstrate that GPT-4-Turbo outperforms crowdworkers in both binary classification (harmful vs. harmless) and multi-label harm categorization tasks. Methodologically, this study extends the application of LLMs to multi-label and multi-modal contexts beyond text annotation and binary classification. Practically, our study contributes to online harm mitigation by guiding the definitions and identification of harmful content on video platforms.

视频安全大模型多模态分类体系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。