arXiv:2601.05232cs.CLcs.CY2026-01

用AI实时分析短视频中的和平度,帮用户了解自己看的视频影响情绪。

AI Application Gives Users Real-Time Feedback on the Level of Peace in the Social Media Videos They Watch

  • 用大语言模型分析视频字幕中的5个和平相关社会维度
  • 模型与人工标注相关性达r²~0.60,显著优于传统情感分析
  • 可为观众和创作者提供媒体内容影响的反馈

如今多数人通过YouTube、Facebook等社交媒体获取新闻,而非传统媒体。内容与语调对冲突或和平有重要影响:仇恨言论加剧对立,和平话语促进和谐。本文开发了一款应用,可实时分析YouTube视频中的语言特征,评估其和平程度。采用两种方法:1)监督学习,基于新闻文本标注的和平水平数据训练全连接前馈网络与两个卷积神经网络,在测试集上准确率达约97%,但在另一新闻数据集仅70%准确,且无法推广至视频转录文本,表明书面语与口语差异显著;2)社会科学研究维度,因缺乏外部标注,使用2种词级情感分析和6种上下文级大语言模型(LLMs)测量59项社会科学研究确认的5个和平维度:同情-轻蔑、新闻-观点、促进-预防、创造-秩序、微妙-简化。结果显示,LLMs在52个视频上的评分与3名人工编码者相关性达r²~0.60,远高于情感分析的r²~0.03。该结果为构建实时反馈系统奠定基础,助力用户与内容创作者评估自身媒体消费与创作的影响。

原文摘要 · Abstract (English)

Most people now get their news from videos on social media, such as YouTube and Facebook, rather than through curated journalism. "We become what we behold." The content and tone of language plays an essential role in starting or ending conflicts. "Hate Speech" can enhance conflict, "Peace Speech" can enhance peace. We developed an application that measures, in real time, these aspects of speech from YouTube videos, which can give users helpful feedback on their own media diet. We used two approaches: 1) supervised machine learning. Language in the text of online news media text was tagged by surveys that measure the level of peace in those countries. One fully connected feedforward and 2 convolutional neural networks trained on that data were $\sim 97\%$ accurate in predicting levels of peace in the test set and $\sim 70\%$ accurate in another distinct news text data set, but did not generalize to YouTube videos, suggesting that written text is different than transcribed spoken language. 2) social science dimensions. There is no similar external data to tag the text in the YouTube video transcripts. We therefore used 2 word-level sentiment analysis (SA) and 6 context-level large language models (LLMs) to measure 5 social dimensions in peace identified by 59 social science studies: compassion-contempt, news-opinion, promotion-prevention, creativity-order, nuance-simplification. LLMs more closely matched the values by 3 human coders on 52 videos, $r^2\sim0.60$ than SA, at $r^2\sim0.03$. Results: LLMs successfully measured social dimensions important in peace in YouTube videos, compared to human coders. These results serve as the basis of an analysis engine that can give users and content creators feedback on their own media diet and creations.

AI伦理视频分析大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。