arXiv:2604.21767cs.CLcs.SI2026-04AAAI

通过音频转录定位视频中谣言片段,提升事实核查精度。

Misinformation Span Detection in Videos via Audio Transcripts

论文配图:Misinformation Span Detection in Videos via Audio Transcripts
图 1 · 摘自论文原文
  • 基于音频转录文本,识别视频中谣言出现的具体时间段。
  • 构建超500个视频、2400+段落的标注数据集,F1达0.68。
  • 开源全部数据与资源,助力谣言检测研究与应用。

在线虚假信息是当前最严峻的挑战之一,可能导致政治极化、民主受创和公共健康风险。虚假信息存在于各类平台及媒体形式中,包括图像、文字、音频和视频。其中,基于视频的虚假信息对事实核查者构成多重挑战,因其易于录制与上传。现有研究多聚焦于视频层面是否传播虚假信息,但无法定位具体谣言发生时段及对应言论。本文构建两个新数据集,通过将视频音频转录为文本,标注虚假信息出现的具体时间段(谣言片段检测)。数据集包含超过500个视频、2400多个标注段落。采用前沿语言模型分类器,实现视频谣言片段检测的F1分数为0.68。所有数据集、转录文本、音频与视频均已公开发布。

原文摘要 · Abstract (English)

Online misinformation is one of the most challenging issues lately, yielding severe consequences, including political polarization, attacks on democracy, and public health risks. Misinformation manifests in any platform with a large user base, including online social networks and messaging apps. It permeates all media and content forms, including images, text, audio, and video. Distinctly, video-based misinformation represents a multifaceted challenge for fact-checkers, given the ease with which individuals can record and upload videos on various video-sharing platforms. Previous research efforts investigated detecting video-based misinformation, focusing on whether a video shares misinformation or not on a video level. While this approach is useful, it only provides a limited and non-easily interpretable view of the problem given that it does not provide an additional context of when misinformation occurs within videos and what content (i.e., claims) are responsible for the video's misinformation nature. In this work, we attempt to bridge this research gap by creating two novel datasets that allow us to explore misinformation detection on videos via audio transcripts, focusing on identifying the span of videos that are responsible for the video's misinformation claim (misinformation span detection). We present two new datasets for this task. We transcribe each video's audio to text, identifying the video segment in which the misinformation claims appears, resulting in two datasets of more than 500 videos with over 2,400 segments containing annotated fact-checked claims. Then, we employ classifiers built with state-of-the-art language models, and our results show that we can identify in which part of a video there is misinformation with an F1 score of 0.68. We make publicly available our annotated datasets. We also release all transcripts, audio and videos.

谣言检测视频分析语音转录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。