arXiv:2409.18346cs.CLcs.CV2024-09被引 13

首个气候议题视频立场检测数据集,助力理解公众观点与传播策略。

MultiClimate: Multimodal Stance Detection on Climate Change Videos

  • 构建100个气候视频+4209帧-字幕对的标注数据集,支持多模态立场分析。
  • 融合图文信息可达到74.7%准确率和74.9%F1,显著优于单模态模型。
  • 发现大模型在多模态立场识别上仍存挑战,适合媒体分析与舆情研究者。

近年来,气候变化(CC)在自然语言处理领域受到越来越多关注。然而,由于缺乏可靠数据集,多模态数据中的立场检测仍被低估且极具挑战。为提升对公众意见和传播策略的理解,本文提出 MultiClimate,首个开源的手动标注立场检测数据集,包含100个与气候相关的YouTube视频及4,209对帧-字幕样本。我们部署了先进的视觉、语言及多模态模型进行立场检测。结果表明,仅使用文本的BERT显著优于仅使用图像的ResNet50和ViT;融合双模态信息达到0.747/0.749的准确率/F1,为当前最佳表现。此外,100M规模的融合模型性能超越CLIP、BLIP,以及更大规模的9B参数IDEFICS和纯文本的Llama3与Gemma2,表明大模型在多模态立场识别任务中仍有较大提升空间。代码、数据及补充材料已公开于https://github.com/werywjw/MultiClimate。

原文摘要 · Abstract (English)

Climate change (CC) has attracted increasing attention in NLP in recent years. However, detecting the stance on CC in multimodal data is understudied and remains challenging due to a lack of reliable datasets. To improve the understanding of public opinions and communication strategies, this paper presents MultiClimate, the first open-source manually-annotated stance detection dataset with $100$ CC-related YouTube videos and $4,209$ frame-transcript pairs. We deploy state-of-the-art vision and language models, as well as multimodal models for MultiClimate stance detection. Results show that text-only BERT significantly outperforms image-only ResNet50 and ViT. Combining both modalities achieves state-of-the-art, $0.747$/$0.749$ in accuracy/F1. Our 100M-sized fusion models also beat CLIP and BLIP, as well as the much larger 9B-sized multimodal IDEFICS and text-only Llama3 and Gemma2, indicating that multimodal stance detection remains challenging for large language models. Our code, dataset, as well as supplementary materials, are available at https://github.com/werywjw/MultiClimate.

多模态立场检测气候议题视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。