用视觉语言模型提升表情包多任务分类性能
On VLMs for Diverse Tasks in Multimodal Meme Classification
- 用VLM解析图像,再微调小LLM理解文本内容
- 在讽刺、攻击性、情感分类上分别提升8.34%、3.52%、26.24%
- 适合研究多模态理解与幽默识别的学者
本文对视觉语言模型(VLMs)在多样表情包分类任务中的表现进行了全面系统分析。提出一种新方法:先用VLM生成对表情包图像的理解,再对小语言模型(LLM)进行文本理解微调,以提升分类性能。贡献包括:(1)针对不同子任务设计多样化提示策略进行VLM基准测试;(2)评估LoRA微调在所有VLM组件上的效果;(3)提出新方案,利用VLM生成的详细图文解释训练小型语言模型,显著提升分类准确率。该联合策略使讽刺、攻击性和情感分类的基线性能分别提升8.34%、3.52%和26.24%。结果揭示了VLM的优劣势,并提供了一种新的表情包理解路径。
原文摘要 · Abstract (English)
In this paper, we present a comprehensive and systematic analysis of vision-language models (VLMs) for disparate meme classification tasks. We introduced a novel approach that generates a VLM-based understanding of meme images and fine-tunes the LLMs on textual understanding of the embedded meme text for improving the performance. Our contributions are threefold: (1) Benchmarking VLMs with diverse prompting strategies purposely to each sub-task; (2) Evaluating LoRA fine-tuning across all VLM components to assess performance gains; and (3) Proposing a novel approach where detailed meme interpretations generated by VLMs are used to train smaller language models (LLMs), significantly improving classification. The strategy of combining VLMs with LLMs improved the baseline performance by 8.34%, 3.52% and 26.24% for sarcasm, offensive and sentiment classification, respectively. Our results reveal the strengths and limitations of VLMs and present a novel strategy for meme understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。