构建多语言视频断言检测数据集,助力识别视频中的可核查言论
ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videos
- 构建跨三语言六主题的视频语料标注集,每句标注可核查/不可核查/观点三类
- 多语言模型在跨验证中达0.896宏F1,但对新领域泛化能力不足
- 为视频谣言检测提供基础工具,适合多模态与信息可信度研究者
视频内容作为传播与虚假信息的重要媒介,其影响力日益增长,亟需有效的分析工具应对多语言、多主题场景下的断言识别。现有反虚假信息研究多聚焦书面文本,忽视了视频转录中口语表达的复杂性。我们提出ViClaim,一个包含1,798个跨英语、德语、西班牙语三种语言、六个主题的视频转录文本标注数据集。每个句子被标注为‘可核查’、‘不可核查’或‘观点’三类。我们开发了定制标注工具以支持复杂的标注流程。实验表明,先进多语言模型在交叉验证中表现优异(宏F1最高达0.896),但在未见主题上的泛化能力仍存挑战,尤其在不同领域间差异显著。研究揭示了视频转录中断言检测的复杂性。ViClaim为推进基于视频的虚假信息检测提供了坚实基础,填补了多模态分析中的关键空白。
原文摘要 · Abstract (English)
The growing influence of video content as a medium for communication and misinformation underscores the urgent need for effective tools to analyze claims in multilingual and multi-topic settings. Existing efforts in misinformation detection largely focus on written text, leaving a significant gap in addressing the complexity of spoken text in video transcripts. We introduce ViClaim, a dataset of 1,798 annotated video transcripts across three languages (English, German, Spanish) and six topics. Each sentence in the transcripts is labeled with three claim-related categories: fact-check-worthy, fact-non-check-worthy, or opinion. We developed a custom annotation tool to facilitate the highly complex annotation process. Experiments with state-of-the-art multilingual language models demonstrate strong performance in cross-validation (macro F1 up to 0.896) but reveal challenges in generalization to unseen topics, particularly for distinct domains. Our findings highlight the complexity of claim detection in video transcripts. ViClaim offers a robust foundation for advancing misinformation detection in video-based communication, addressing a critical gap in multimodal analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。