构建中文多面板梗图数据集,检验模型是否真懂画面顺序。
Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning

- 设计1214个带顺序约束的中文梗图样本,涵盖五种结构类型。
- 模型在乱序画面下准确率骤降,暴露对顺序理解不足。
- 适合研究视觉语言推理、跨模态顺序感知的研究者使用。
许多多模态任务依赖于视觉元素的排列与组合方式,而不仅仅是孤立识别。网络梗图是这一问题的紧凑体现:其笑点往往依赖于特定阅读顺序及跨面板的图文线索。尽管大视觉语言模型(LVLMs)在单图理解上表现强劲,但它们在中文社交媒体中对结构化梗图布局的序列感知推理能力仍不明确。本文引入CMPM,一个包含1,214个标注样本的中文多面板梗图基准,覆盖五种结构类型、顺序依赖性、面板顺序约束及可选评论上下文。我们设计两层评估:Task1测试结构分类与顺序敏感的面板排序(含上下文消融设置);Task2通过人类评分在五个1-3级李克特量表(视觉、面板、幽默、上下文、忠实度)上评估中文梗图解释生成。在统一协议下对五种代表性LVLM进行基准测试。结果显示,标准显示准确率本身不能证明顺序理解能力:主乱序条件下准确率显著下降,揭示了顺序敏感多模态推理的持续差距。Task2偏好表明Gemini 3.1 Pro和GPT-5.5优于开源模型,而评论上下文仅带来微小且混合的提升。代码与数据将在接受后发布。
原文摘要 · Abstract (English)
Many multimodal tasks depend on how visual elements are ordered and composed, not only on recognizing them in isolation. Internet memes are a compact case of this problem: their punchline often depends on a constrained reading order and cross-panel visual--textual cues. While large vision-language models (LVLMs) show strong performance on single-image understanding, it remains unclear whether they can perform sequence-aware reasoning over structured meme layouts, especially in Chinese social media. We introduce CMPM, a Chinese Multi-Panel Meme benchmark with 1,214 annotated samples covering five structural types, ordering dependency, panel-order constraints, and optional comment context. We formulate a two-layer evaluation: Task1 probes structure typing and order-sensitive panel sequencing (with a context ablation setting), and Task2 evaluates Chinese meme explanation generation with human ratings on five 1-3 Likert dimensions (visual, panel, humor, context, and faithfulness). We benchmark five representative LVLMs under a unified protocol. Results indicate that canonical-display accuracy is not by itself evidence of order understanding: the primary shuffled condition produces a sharp accuracy drop, revealing a persistent gap in order-sensitive multimodal reasoning. Task2 preferences place Gemini 3.1 Pro and GPT-5.5 above the open models, while comment context yields only a small and mixed Core4 gain. Code and data will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。