arXiv:2606.14780cs.CVcs.LG2026-06

21K条视频构建多模态点击诱饵数据集,助力平台内容审核。

YTClickbait21K: Human-Annotated Multimodal Dataset for YouTube Clickbait Detection Across Diverse Channels and Content Categories

  • 从40个频道收集21,238个视频,含标题、描述、缩略图等多模态信息
  • 三名标注员独立打标,多数投票达成一致,一致性k=0.65
  • 适用于跨模态理解与自动化内容审核研究,覆盖新闻、娱乐等多领域

视频分享平台上的点击诱饵内容严重威胁信息可信度,但自动化检测因缺乏大规模高质量多模态数据集而受限。我们提出YTClickbait21K,一个包含21,238个视频的人工标注YouTube点击诱饵数据集,涵盖来自29个国家的40个频道,覆盖新闻、娱乐、教育和游戏等多种内容类别。每个样本包含结构化元数据(标题、描述、互动统计数据)及对应的缩略图,支持全面的多模态分析。为保证标注质量,每条视频由三位标注员独立使用标准化决策框架标注,结合文本、视觉及跨模态一致性线索,最终标签通过多数投票确定。数据集展现出较高的标注者间一致性(k=0.65),证实了标注的可靠性,尽管点击诱饵判定本身具有主观性。该数据集兼具规模、标注严谨性和多模态丰富性,为机器学习模型开发与评估提供了可靠基准,推动跨模态语义理解研究,并促进自动化内容审核系统的发展。

原文摘要 · Abstract (English)

Clickbait content on video-sharing platforms poses a significant challenge to information reliability, yet progress in automated detection has been constrained by the lack of large-scale, high-quality multimodal datasets. We present YTClickbait21K, a human-annotated YouTube clickbait dataset comprising 21,238 videos collected from 40 channels across 29 countries, covering diverse content categories such as news, entertainment, education, and gaming. Each sample includes structured metadata (title, description, engagement statistics) along with associated thumbnail images, enabling comprehensive multimodal analysis. To ensure annotation quality, every video was independently labeled by three annotators using a standardized decision framework that incorporates textual, visual, and cross-modal consistency cues, with final labels determined through majority voting. The dataset exhibits substantial inter-annotator agreement (k=0.65), confirming reliable labeling despite the inherent subjectivity of clickbait detection. By combining scale, annotation rigor, and multimodal richness, this dataset provides a robust benchmark for developing and evaluating machine learning models, facilitating research in cross-modal semantic understanding, and advancing automated content moderation systems.

点击诱饵多模态数据集内容审核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。