arXiv:2509.08024cs.CVcs.CY2025-09被引 2

用大模型融合图文信息,提升气候议题立场判断准确率。

Two Stage Context Learning with Large Language Models for Multimodal Stance Detection on Climate Change

  • 分两阶段提取文本摘要与视觉描述,再联合建模
  • 在MultiClimate数据集上达76.2%的F1分数
  • 适合关注多模态舆情分析的研究者

随着数字平台信息爆炸式增长,立场检测成为社交媒体分析的关键挑战。现有方法多仅依赖文本,但真实社交内容常含图文混合信息,亟需先进多模态方法。为此,我们提出一种多模态立场检测框架,通过分层融合策略整合文本与视觉信息。首先利用大语言模型从源文本中检索与立场相关摘要,同时采用领域感知的图像描述生成器,结合目标主题解析视觉内容。随后,将这些模态与回复文本共同输入专用Transformer模块,捕捉图文间交互关系。该模态融合框架有效提升立场分类鲁棒性。我们在包含对齐视频帧与字幕的MultiClimate基准数据集上评估,取得76.2%准确率、76.3%精确率、76.2%召回率和76.2% F1分数,优于现有最先进方法。

原文摘要 · Abstract (English)

With the rapid proliferation of information across digital platforms, stance detection has emerged as a pivotal challenge in social media analysis. While most of the existing approaches focus solely on textual data, real-world social media content increasingly combines text with visual elements creating a need for advanced multimodal methods. To address this gap, we propose a multimodal stance detection framework that integrates textual and visual information through a hierarchical fusion approach. Our method first employs a Large Language Model to retrieve stance-relevant summaries from source text, while a domain-aware image caption generator interprets visual content in the context of the target topic. These modalities are then jointly modeled along with the reply text, through a specialized transformer module that captures interactions between the texts and images. The proposed modality fusion framework integrates diverse modalities to facilitate robust stance classification. We evaluate our approach on the MultiClimate dataset, a benchmark for climate change-related stance detection containing aligned video frames and transcripts. We achieve accuracy of 76.2%, precision of 76.3%, recall of 76.2% and F1-score of 76.2%, respectively, outperforming existing state-of-the-art approaches.

多模态立场检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。