评测社交媒体回复对原文的多模态立场理解能力,发现大模型仍难把握上下文关系。
MMDS-Bench: Benchmarking Multimodal Large Language Models on Dynamic Stance in Social Media Interactions

- 构建包含3482条数据的多模态立场诊断基准,覆盖图文混合场景。
- 12个大模型在需推理关系的任务上表现不佳,准确率低于60%。
- 适合研究多模态对话理解、社交文本分析的学者和工程师参考。
动态立场分类关注回复如何回应其直接父消息,而非帖子与固定话题的关系。现有研究主要聚焦纯文本场景,而社交媒体互动日益依赖图像、截图、表情包、反应图及跨模态引用。本文提出MMDS-Bench,一个针对社交媒体父子回复交互中多模态动态立场分类的诊断基准。该基准包含3,482个多模态样本,标注七类动态立场,并提供800个需要结构化推理的诊断子集,涵盖父消息理解、回复理解与立场关系推断。每条样本还标注五类挑战因素:多模态融合、父消息框架、非字面表达、互动推理和标签边界模糊。我们评估了12个闭源与开源多模态大语言模型,并提出基于参考的LLM评判协议以评估推理质量。结果表明,当前多模态大模型在需要超越独立理解父回复的关联推理任务中仍存在显著困难。
原文摘要 · Abstract (English)
Dynamic stance classification models how a reply responds to its direct parent message, rather than how a post relates to a fixed topic. Existing work has mainly studied this problem in text-only settings, while social media interactions increasingly rely on images, screenshots, memes, reaction images, and cross-modal references. We introduce MMDS-Bench, a diagnostic benchmark for multimodal dynamic stance classification in social media parent-reply interactions. MMDS-Bench contains 3,482 multimodal instances annotated with a seven-label dynamic stance taxonomy, together with an 800-instance diagnostic subset that requires structured reasoning over parent understanding, reply understanding, and stance-relation inference. We further annotate each instance with five challenge factors covering multimodal fusion, parent framing, non-literal expression, interaction reasoning, and label-boundary ambiguity. We evaluate 12 closed-source and open-source multimodal large language models and propose a reference-grounded LLM-judge protocol for assessing reasoning quality. Results show that current MLLMs still struggle with multimodal dynamic stance understanding, especially in cases that require relational inference beyond separate parent and reply comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。