arXiv:2501.11770cs.CLcs.CY2025-01被引 9

从TikTok短视频中提取博主隐含的价值观,揭示青少年价值观传播新路径。

The Value of Nothing: Multimodal Extraction of Human Values Expressed by TikTok Influencers

  • 分两步处理:先转成文字脚本,再用大模型识别价值观
  • 两步法比直接识别效果提升显著,少样本LLM优于微调MLM
  • 首次公开标注了数百个TikTok视频的价值观数据集

社会与个人价值观通过互动和接触传递给年轻一代。传统上,儿童和青少年通过父母、教师或同龄人学习价值观。如今,社交媒体成为青年(及成人)获取信息的主要渠道,既是娱乐媒介,也可能成为价值观习得的途径。本文从针对儿童和青少年的TikTok网红视频中提取隐含价值观。我们构建了一个包含数百个视频的数据集,并依据舒瓦茨个人价值理论进行标注。实验比较了多种语言模型在价值观识别中的表现,采用两种流程:直接从视频提取,以及先将视频转换为详尽脚本,再从文本中提取。结果表明,两步法显著优于直接法;在两个阶段均使用少样本大语言模型的表现优于在第二阶段使用微调掩码语言模型。我们还探讨了持续预训练与微调的影响,并对比不同模型在识别视频中支持或反对的价值观上的表现。最后,我们发布了首个经过价值观标注的TikTok视频数据集。据我们所知,这是首次专门针对TikTok乃至视觉社交媒体的价值提取尝试。研究结果为未来基于视频社交平台的价值传播研究奠定了基础。

原文摘要 · Abstract (English)

Societal and personal values are transmitted to younger generations through interaction and exposure. Traditionally, children and adolescents learned values from parents, educators, or peers. Nowadays, social platforms serve as a significant channel through which youth (and adults) consume information, as the main medium of entertainment, and possibly the medium through which they learn different values. In this paper we extract implicit values from TikTok movies uploaded by online influencers targeting children and adolescents. We curated a dataset of hundreds of TikTok movies and annotated them according to the well established Schwartz Theory of Personal Values. We then experimented with an array of language models, investigating their utility in value identification. Specifically, we considered two pipelines: direct extraction of values from video and a 2-step approach in which videos are first converted to elaborated scripts and values are extracted from the textual scripts. We find that the 2-step approach performs significantly better than the direct approach and that using a few-shot application of a Large Language Model in both stages outperformed the use of a fine-tuned Masked Language Model in the second stage. We further discuss the impact of continuous pretraining and fine-tuning and compare the performance of the different models on identification of values endorsed or confronted in the TikTok. Finally, we share the first values-annotated dataset of TikTok videos. To the best of our knowledge, this is the first attempt to extract values from TikTok specifically, and visual social media in general. Our results pave the way to future research on value transmission in video-based social platforms.

价值观识别TikTok多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。