arXiv:2409.01037cs.CL2024-09

构建首个图文隐喻与反语理解基准,助力模型读懂网络梗的深层含义。

NYK-MS: A Well-annotated Multi-modal Metaphor and Sarcasm Understanding Benchmark on Cartoon-Caption Dataset

  • 构建含1583个隐喻和1578个反语样本的图文标注数据集
  • 多轮标注+GUI与GPT-4V辅助,确保7项任务标注一致性
  • 验证大模型在零样本下理解图文讽刺能力的不足

隐喻与反语是网络交流中常见的修辞表达,尤其在青少年流行的迷因(meme)中广泛出现。本文构建了一个名为NYK-MS(NewYorKer for Metaphor and Sarcasm)的新基准,包含1,583个用于隐喻理解的任务样本和1,578个用于反语理解的任务样本。该数据集涵盖是否含隐喻/反语、具体哪一词或物体承载隐喻/反语、讽刺对象及原因等7项任务,均由至少3名标注者进行精细标注。通过多轮标注与人工校验提升质量,并借助GUI界面与GPT-4V提升效率。基于此基准,我们进行了大量实验:零样本测试表明,大型语言模型(LLM)与大型多模态模型(LMM)在分类任务上表现不佳,但随着模型规模增大,其余5项任务性能有所提升;在传统预训练模型上的实验则表明,增强与对齐方法能有效提升性能,证明该基准与现有数据集具有一致性,且要求模型具备跨模态理解能力。

原文摘要 · Abstract (English)

Metaphor and sarcasm are common figurative expressions in people's communication, especially on the Internet or the memes popular among teenagers. We create a new benchmark named NYK-MS (NewYorKer for Metaphor and Sarcasm), which contains 1,583 samples for metaphor understanding tasks and 1,578 samples for sarcasm understanding tasks. These tasks include whether it contains metaphor/sarcasm, which word or object contains metaphor/sarcasm, what does it satirize and why does it contains metaphor/sarcasm, all of the 7 tasks are well-annotated by at least 3 annotators. We annotate the dataset for several rounds to improve the consistency and quality, and use GUI and GPT-4V to raise our efficiency. Based on the benchmark, we conduct plenty of experiments. In the zero-shot experiments, we show that Large Language Models (LLM) and Large Multi-modal Models (LMM) can't do classification task well, and as the scale increases, the performance on other 5 tasks improves. In the experiments on traditional pre-train models, we show the enhancement with augment and alignment methods, which prove our benchmark is consistent with previous dataset and requires the model to understand both of the two modalities.

隐喻理解反语识别图文理解多模态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。