检测儿童向YouTube视频中的商业广告,识别推广内容与风险
ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos
- 基于3360个视频构建评测任务,分层提供字幕、链接等证据
- 超45%视频未正确使用平台广告标注功能,存在合规风险
- 使用GPT-5系列模型生成标签,适合内容安全与合规研究者
ChildSafeAds是针对可能触达儿童和青少年的YouTube视频中商业内容的共享任务,包含来自939个频道的3360个视频实例。每个实例起始于SponsorBlock用户标记的赞助片段,附带字幕、视频及频道信息,以及视频描述中链接的销售或服务页面。参赛系统需完成三项任务:判断推广类型(ST1)、分配产品类别(ST2)、识别法律风险标志(ST3)。证据按四层累积访问权限提供,从字幕到链接页面,便于评估数据获取成本。数据显示,45.5%的视频未正确使用平台内广告披露标识(“包含付费推广”标签)。GPT-5.4在专家团队审阅样本并迭代分类体系、提示词与模型选择后生成标签;GPT-5.6-luna独立标注开发集。本文介绍任务设计、数据与评估方案,更新版本将补充参与系统与共享任务结果。
原文摘要 · Abstract (English)
ChildSafeAds is a shared task on commercial content in YouTube videos likely to reach children and teenagers. It contains 3,360 videos from 939 channels. Each instance begins with a segment submitted to SponsorBlock, an open-source crowdsourced browser extension whose users mark sponsor segments so that others can skip them. We pair the segment with its available transcript, video and channel information, and a sales or service page linked from the video description. Systems determine what kind of offer is being promoted (ST1), assign product categories (ST2), and identify legal risk flags (ST3). The evidence is divided into four cumulative access levels, from the transcript to the linked page, so results can be compared against the cost of collecting the data. 45.5\% of videos in our data failed to properly use the in-platform ad disclosure method (the ``Includes paid promotion'' label). GPT-5.4 produced the labels after the expert organiser team reviewed samples and iterated on the taxonomy, prompts and model choices. GPT-5.6-luna independently labelled the development set. This report describes the task, data and evaluation. An updated version will add participating systems and shared-task results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。