arXiv:2607.24191cs.CLcs.AI2026-07

构建多模态对话立场反转预测基准,提升对情绪与逻辑交织的动态理解能力。

StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

论文配图:StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting
图 1 · 摘自论文原文
  • 提出多维度立场提取与动态反转追踪双任务,捕捉对话中立场演变
  • 在五个模态上实现领先性能,尤其在立场反转触发因素识别上显著超越基线
  • 适合研究多模态对话理解、情感推理与大模型可解释性的学者使用

对话立场检测已从静态文本分析转向动态多模态建模。然而现有基准存在三大缺陷:难以捕捉信念动态演变,特别是立场反转过程;难以区分情感状态与逻辑推理;忽视多模态线索在化解语用歧义(如反讽)中的关键作用。为此,我们提出StanceFlip,一个面向多轮对话中跨五模态、多场景的多模态对话立场反转预测基准,包含两个新子任务:1)多模态立场六元组提取,从对话中提取持有者、目标、情绪、情感、立场与理由,作为认知结构的静态快照;2)动态立场反转归因,追踪对话中的立场转变并识别其底层触发因素。同时提出专用于多模态对话立场反转预测(MCSFF)的ConStaFF框架。基于大语言模型,ConStaFF采用思想-立场(ToS)推理框架与自我反思验证机制,实现端到端立场推理。ToS将推理过程分解为多个认知角色,分别处理目标命题构建、跨模态冲突解决与历史立场轨迹推断。大量实验表明,该方法在六元组提取与反转触发归因任务上均达到当前最优表现,显著优于强大多模态大模型基线。

原文摘要 · Abstract (English)

Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly during stance reversals; difficulty in disentangling affective states from logical reasoning; and neglect of the critical role of multimodal cues in resolving pragmatic ambiguities such as sarcasm. To address these limitations, we propose StanceFlip, a benchmark designed for multimodal conversational stance flipping forecasting over multi-turn dialogues across five modalities and multi-scenarios, which includes two novel subtasks: 1) Multimodal Stance Sextuple Extraction, extracting holder, target, emotion, sentiment, stance, and rationale as static state snapshots of dialogue to capture fine-grained cognitive structures. 2) Dynamic Stance Flip Attribution, tracking stance reversals across the conversation and identifying their underlying triggers. Alongside the dataset, we propose a dedicated framework, named ConStaFF, for Multimodal Conversational Stance Flipping Forecasting (MCSFF). Built upon a large language model, ConStaFF performs end-to-end stance reasoning, with a Thought-of-Stance (ToS) reasoning framework and a self-reflective verification mechanism integrated for structured stance modeling and faithful flip attribution. Specifically, ToS decomposes the reasoning process into specialized cognitive personas to formulate target propositions, resolve cross-modal conflicts, and infer historical stance trajectories. Extensive experiments show that our approach achieves state-of-the-art performance on both sextuple extraction and flip-trigger attribution, outperforming strong multimodal large language model baselines by substantial margins.

多模态对话立场检测大模型推理反讽识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。