arXiv:2504.11108cs.CL2025-04中稿 · ed被引 4

评测视觉语言模型在德语事实知识上的表现,发现其对德语图像理解存在明显短板。

Benchmarking Vision Language Models on German Factual Data

  • 通过双语提示与德语/国际图像对比,分离视觉与文本能力差异
  • 名人和景点识别准确率低,因缺乏对德语图像内容的视觉认知
  • 动植物识别依赖英文名,德语表达仍不准确,适合多语言研究者参考

与大语言模型类似,视觉语言模型的发展主要依赖英语数据集和英语、中文训练模型,而对其他语言的支持,即使像德语这样高资源语言,也显著不足。本文分析了开源视觉语言模型在德语与英语事实知识上的表现。通过‘裁判即评委’的方式,在德语与国际语境的图像中,分别测试模型在两种语言提示下的准确率,以区分视觉与文本能力。结果显示:对于名人和景点,模型表现较差,因缺乏对德语图像内容的视觉认知;对于动物和植物,模型能根据科学名或英文通用名正确识别图像内容,但在德语表达上失败;汽车与超市商品在英语和德语图像中识别表现相当,且在两种提示语言下均表现良好。

原文摘要 · Abstract (English)

Similar to LLMs, the development of vision language models is mainly driven by English datasets and models trained in English and Chinese language, whereas support for other languages, even those considered high-resource languages such as German, remains significantly weaker. In this work we present an analysis of open-weight VLMs on factual knowledge in the German and English language. We disentangle the image-related aspects from the textual ones by analyzing accu-racy with jury-as-a-judge in both prompt languages and images from German and international contexts. We found that for celebrities and sights, VLMs struggle because they are lacking visual cognition of German image contents. For animals and plants, the tested models can often correctly identify the image contents ac-cording to the scientific name or English common name but fail in German lan-guage. Cars and supermarket products were identified equally well in English and German images across both prompt languages.

视觉语言模型多语言德语评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。