arXiv:2508.06515cs.CV2025-08Conference of the …被引 2

用临床专家标注数据,识别TikTok上鼓吹肌肉畸形的有害内容。

BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok

  • 基于临床医生标注构建多模态数据集,精准识别有害健身内容。
  • 视频特征比文本更有效,多模态融合提升检测准确率5%~15%。
  • 适合平台安全团队与心理健康研究者参考,推动智能内容审核。

社交媒体平台在识别助长肌肉畸形行为(bigorexia)的有害内容方面面临严峻挑战,此类内容常伪装成合法健身建议,主要影响青少年男性。本文提出BigTokDetect框架,针对TikTok上的此类内容进行检测。我们构建了BigTok数据集,包含超过2,200条由临床精神病学家标注的TikTok视频,涵盖五类主类别和十八个细粒度子类别。对现有先进视觉-语言模型的全面评估表明,尽管商业零样本模型在主类别上表现最佳,但经过微调的小型开源模型在细粒度子类别检测中表现更优。消融实验显示,多模态融合使性能提升5%至15%,其中视频特征提供最具区分性的信号。研究支持一种基于实证的审核策略:自动识别明确危害内容,将模糊内容标记供人工复核,并为新兴心理卫生领域建立可扩展的防护框架。

原文摘要 · Abstract (English)

Social media platforms face escalating challenges in detecting harmful content that promotes muscle dysmorphic behaviors and cognitions (bigorexia). This content can evade moderation by camouflaging as legitimate fitness advice and disproportionately affects adolescent males. We address this challenge with BigTokDetect, a clinically informed framework for identifying pro-bigorexia content on TikTok. We introduce BigTok, the first expert-annotated multimodal benchmark dataset of over 2,200 TikTok videos labeled by clinical psychiatrists across five categories and eighteen fine-grained subcategories. Comprehensive evaluation of state-of-the-art vision-language models reveals that while commercial zero-shot models achieve the highest accuracy on broad primary categories, supervised fine-tuning enables smaller open-source models to perform better on fine-grained subcategory detection. Ablation studies show that multimodal fusion improves performance by 5 to 15 percent, with video features providing the most discriminative signals. These findings support a grounded moderation approach that automates detection of explicit harms while flagging ambiguous content for human review, and they establish a scalable framework for harm mitigation in emerging mental health domains.

内容检测视觉语言模型心理健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。