arXiv:2603.23627cs.CVcs.AI2026-03被引 6

构建乌克兰语视觉词义消歧基准,揭示多语言模型性能差距

Ukrainian Visual Word Sense Disambiguation Benchmark

  • 基于类比英/意/波斯语方法,半自动构建乌克兰语视觉词义消歧数据集
  • 八款多模态大模型均低于英语零样本CLIP基线,表现显著更差
  • 揭示乌语与英语在视觉词义消歧任务中的巨大性能鸿沟,适合跨语言研究

本研究提出了乌克兰语视觉词义消歧(Visual-WSD)任务的基准。该任务旨在仅依赖极少上下文信息,从一组十张图像中识别出目标歧义词最恰当的语义表示。为构建该基准,我们采用与先前英、意、波斯语研究相似的方法论,使乌克兰语基准可融入跨语言模型性能对比框架。数据通过半自动方式收集,并经领域专家修正。我们评估了八种多语言及多模态大语言模型的表现。所有模型均不及(引用文献)用于英语Visual-WSD任务的零样本CLIP基线模型。分析显示,乌克兰语与英语在Visual-WSD任务上存在显著性能差距。

原文摘要 · Abstract (English)

This study presents a benchmark for evaluating the Visual Word Sense Disambiguation (Visual-WSD) task in Ukrainian. The main goal of the Visual-WSD task is to identify, with minimal contextual information, the most appropriate representation of a given ambiguous word from a set of ten images. To construct this benchmark, we followed a methodology similar to that proposed by (CITATION), who previously introduced benchmarks for the Visual-WSD task in English, Italian, and Farsi. This approach allows us to incorporate the Ukrainian benchmark into a broader framework for cross-language model performance comparisons. We collected the benchmark data semi-automatically and refined it with input from domain experts. We then assessed eight multilingual and multimodal large language models using this benchmark. All tested models performed worse than the zero-shot CLIP-based baseline model (CITATION) used by (CITATION) for the English Visual-WSD task. Our analysis revealed a significant performance gap in the Visual-WSD task between Ukrainian and English.

视觉词义消歧多语言基准测试乌克兰语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。