评测视觉语言模型在多模态多语言立场检测中的表现
Exploring Vision Language Models for Multimodal and Multilingual Stance Detection
- 用多语言图文数据集测试先进视觉语言模型
- 模型主要依赖文本,尤其图像中的文字
- 跨语言表现稳定,但个别模型异常
社交媒体的全球传播加剧了信息扩散,凸显了跨语言与多模态自然语言处理任务(如立场检测)的重要性。以往研究多集中于纯文本输入,对图文并存场景关注不足。近年来,多模态内容占比显著上升。尽管当前先进的视觉语言模型(VLMs)展现出潜力,其在多模态多语言立场检测任务上的表现仍缺乏系统评估。本文在新扩展的数据集上评测了多个前沿VLMs,涵盖七种语言和多模态输入,分析其对视觉线索的利用、语言特异性表现及跨模态交互。结果表明,VLMs普遍更依赖文本而非图像,且该趋势在各语言中保持一致;尤其对图像中的文字内容依赖更强。在多语言方面,模型预测在多数情况下跨语言一致,无论是否为显式多语言模型,但存在少数与宏平均F1、语言支持范围或模型规模不匹配的异常情况。
原文摘要 · Abstract (English)
Social media's global reach amplifies the spread of information, highlighting the need for robust Natural Language Processing tasks like stance detection across languages and modalities. Prior research predominantly focuses on text-only inputs, leaving multimodal scenarios, such as those involving both images and text, relatively underexplored. Meanwhile, the prevalence of multimodal posts has increased significantly in recent years. Although state-of-the-art Vision-Language Models (VLMs) show promise, their performance on multimodal and multilingual stance detection tasks remains largely unexamined. This paper evaluates state-of-the-art VLMs on a newly extended dataset covering seven languages and multimodal inputs, investigating their use of visual cues, language-specific performance, and cross-modality interactions. Our results show that VLMs generally rely more on text than images for stance detection and this trend persists across languages. Additionally, VLMs rely significantly more on text contained within the images than other visual content. Regarding multilinguality, the models studied tend to generate consistent predictions across languages whether they are explicitly multilingual or not, although there are outliers that are incongruous with macro F1, language support, and model size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。