arXiv:2409.00597cs.MMcs.CL2024-09被引 30

构建首个多轮对话场景下的多模态立场检测数据集与模型。

Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model

  • 提出联合文本与视觉模态的多模态大模型框架,捕捉对话中立场演变。
  • 在新数据集上达到当前最优性能,准确率显著超越基线方法。
  • 适合研究社交网络舆情分析、多模态信息理解的学者与工程师。

立场检测旨在利用社交媒体数据识别公众对特定目标的态度,是重要但具有挑战性的任务。随着包含文本和图像的多样化多模态社交媒体内容激增,多模态立场检测(MSD)成为关键研究方向。然而,现有研究仅关注单个图文对中的立场建模,忽视了社交媒体中自然存在的多方对话上下文。这一局限源于缺乏真实反映此类对话场景的数据集,制约了对话式多模态立场检测的发展。为此,我们提出了一个全新的多模态多轮对话立场检测数据集(MmMtCSD)。为从该挑战性数据集中推断立场,我们设计了一种新型多模态大语言模型立场检测框架(MLLM-SD),通过联合学习文本与视觉模态的立场表示。在MmMtCSD上的实验表明,所提的MLLM-SD方法在多模态立场检测任务中达到最先进的性能。我们相信,MmMtCSD将推动立场检测研究在现实应用中的发展。

原文摘要 · Abstract (English)

Stance detection, which aims to identify public opinion towards specific targets using social media data, is an important yet challenging task. With the proliferation of diverse multimodal social media content including text, and images multimodal stance detection (MSD) has become a crucial research area. However, existing MSD studies have focused on modeling stance within individual text-image pairs, overlooking the multi-party conversational contexts that naturally occur on social media. This limitation stems from a lack of datasets that authentically capture such conversational scenarios, hindering progress in conversational MSD. To address this, we introduce a new multimodal multi-turn conversational stance detection dataset (called MmMtCSD). To derive stances from this challenging dataset, we propose a novel multimodal large language model stance detection framework (MLLM-SD), that learns joint stance representations from textual and visual modalities. Experiments on MmMtCSD show state-of-the-art performance of our proposed MLLM-SD approach for multimodal stance detection. We believe that MmMtCSD will contribute to advancing real-world applications of stance detection research.

立场检测多模态对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。