多智能体框架融合检索增强,提升图文立场判断准确性
MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection

- 引入检索增强与多智能体协作,分步处理图文信息
- 在五个数据集上超越现有方法,最高提升4.7%准确率
- 适合需要细粒度跨模态分析的舆情研究者
多模态立场检测对理解公众舆论至关重要,但有效融合文本与图像(尤其存在矛盾信号时)仍具挑战。现有方法常面临上下文定位不准、跨模态解释模糊及单次推理脆弱等问题。为此,我们提出检索增强型多模态多智能体立场检测框架(MM-StanceDet),集成检索增强实现上下文定位,专用多模态分析智能体进行细致解读,推理增强辩论阶段探索多元视角,自省机制实现稳健判定。在五个数据集上的大量实验表明,MM-StanceDet显著优于当前最优基线,验证了其多智能体架构与结构化推理流程在应对复杂多模态立场任务中的有效性。
原文摘要 · Abstract (English)
Multimodal Stance Detection (MSD) is crucial for understanding public discourse, yet effectively fusing text and image, especially with conflicting signals, remains challenging. Existing methods often face difficulties with contextual grounding, cross-modal interpretation ambiguity, and single-pass reasoning fragility. To address these, we propose Retrieval-Augmented Multi-modal Multi-agent Stance Detection (MM-StanceDet), a novel multi-agent framework integrating Retrieval Augmentation for contextual grounding, specialized Multimodal Analysis agents for nuanced interpretation, a Reasoning-Enhanced Debate stage for exploring perspectives, and Self-Reflection for robust adjudication. Extensive experiments on five datasets demonstrate MM-StanceDet significantly outperforms state-of-the-art baselines, validating the efficacy of its multi-agent architecture and structured reasoning stages in addressing complex multimodal stance challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。