arXiv:2505.05040cs.CLcs.AI2025-05

研究多语言推文图文关系,构建了拉脱维亚语-英语双语数据集。

Image-Text Relation Prediction for Multilingual Tweets

  • 用推文图文对构建多语言关系预测数据集
  • 新模型在多语言图文匹配上表现优于旧模型
  • 适合做跨语言视觉语言理解的研究者参考

社交媒体平台允许用户上传图文已逾十年,但图文间关系仍不明确。本文研究多语言视觉语言模型在不同语言中进行图文关系预测的能力,从推文中构建了一个包含拉脱维亚语及其人工翻译成英语的平衡基准数据集。通过对比现有工作,发现越近期发布的视觉语言模型在该任务上表现越强,但仍存在显著提升空间。

原文摘要 · Abstract (English)

Various social networks have been allowing media uploads for over a decade now. Still, it has not always been clear what is their relation with the posted text or even if there is any at all. In this work, we explore how multilingual vision-language models tackle the task of image-text relation prediction in different languages, and construct a dedicated balanced benchmark data set from Twitter posts in Latvian along with their manual translations into English. We compare our results to previous work and show that the more recently released vision-language model checkpoints are becoming increasingly capable at this task, but there is still much room for further improvement.

图文关系多语言社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。