提出双推理框架MIND,让多模态立场检测更懂人类的直觉与反思。
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
- 用直觉-反思双过程模拟人类推理,动态更新模态与语义经验池。
- 在MMSD数据集上准确率超多数基线模型,泛化能力显著提升。
- 适合研究多模态理解、人机认知对齐及社交舆情分析的学者。
多模态立场检测(MSD)是理解社交媒体公众意见的关键任务。现有方法主要聚焦于模态融合,缺乏显式推理机制来捕捉模态间动态(如反讽或冲突)如何共同影响用户最终立场,常导致误判。为此,我们倡导从‘学习融合’转向‘学习推理’。提出MIND:一种元认知直觉-反思网络,用于双推理。受人类认知双过程理论启发,MIND构建自优化循环:先通过动态更新的模态与语义经验池快速生成直觉假设;再由元认知反思阶段利用模态-思维链(Modality-CoT)和语义-思维链(Semantic-CoT)检验初始判断,提炼优策略并进化经验池。双经验结构在训练中持续精炼,并在推理时调用,实现鲁棒、上下文感知的立场决策。在MMSD基准上的大量实验表明,MIND显著优于多数基线模型,展现出强泛化性。
原文摘要 · Abstract (English)
Multimodal Stance Detection (MSD) is a crucial task for understanding public opinion on social media. Existing methods predominantly operate by learning to fuse modalities. They lack an explicit reasoning process to discern how inter-modal dynamics, such as irony or conflict, collectively shape the user's final stance, leading to frequent misjudgments. To address this, we advocate for a paradigm shift from *learning to fuse* to *learning to reason*. We introduce **MIND**, a **M**eta-cognitive **I**ntuitive-reflective **N**etwork for **D**ual-reasoning. Inspired by the dual-process theory of human cognition, MIND operationalizes a self-improving loop. It first generates a rapid, intuitive hypothesis by querying evolving Modality and Semantic Experience Pools. Subsequently, a meta-cognitive reflective stage uses Modality-CoT and Semantic-CoT to scrutinize this initial judgment, distill superior adaptive strategies, and evolve the experience pools themselves. These dual experience structures are continuously refined during training and recalled at inference to guide robust and context-aware stance decisions. Extensive experiments on the MMSD benchmark demonstrate that our MIND significantly outperforms most baseline models and exhibits strong generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。