用视觉信息解决词语歧义,让AI更懂语言与图像的关联。
Bridging Lexical Ambiguity and Vision: A Mini Review on Visual Word Sense Disambiguation
- 结合图像与文本,通过对比学习等方法识别词义。
- 基于CLIP和大模型的系统在准确率上提升6-8%。
- 适合多模态理解、跨语言应用的研究者参考。
本文综述了视觉词义消歧(VWSD),即传统词义消歧(WSD)在多模态场景下的延伸。传统WSD仅依赖文本和词典资源,而VWSD利用视觉线索,在极少文本输入下确定歧义词的正确含义。文章回顾了2016至2025年间的发展,涵盖基于特征、图结构及对比嵌入的技术演进,重点分析提示工程、微调策略与多语言适配。量化结果表明,经CLIP微调的模型与融合大语言模型(LLM)的系统显著优于零样本基线,均在平均倒数排名(MRR)上提升6-8%。然而仍面临上下文有限、模型偏好常见词义、缺乏多语言数据集及评估框架不完善等挑战。未来方向在于整合CLIP对齐、扩散生成与大模型推理能力,构建更具上下文感知与多语言支持的强健系统。
原文摘要 · Abstract (English)
This paper offers a mini review of Visual Word Sense Disambiguation (VWSD), which is a multimodal extension of traditional Word Sense Disambiguation (WSD). VWSD helps tackle lexical ambiguity in vision-language tasks. While conventional WSD depends only on text and lexical resources, VWSD uses visual cues to find the right meaning of ambiguous words with minimal text input. The review looks at developments from early multimodal fusion methods to new frameworks that use contrastive models like CLIP, diffusion-based text-to-image generation, and large language model (LLM) support. Studies from 2016 to 2025 are examined to show the growth of VWSD through feature-based, graph-based, and contrastive embedding techniques. It focuses on prompt engineering, fine-tuning, and adapting to multiple languages. Quantitative results show that CLIP-based fine-tuned models and LLM-enhanced VWSD systems consistently perform better than zero-shot baselines, achieving gains of up to 6-8\% in Mean Reciprocal Rank (MRR). However, challenges still exist, such as limitations in context, model bias toward common meanings, a lack of multilingual datasets, and the need for better evaluation frameworks. The analysis highlights the growing overlap of CLIP alignment, diffusion generation, and LLM reasoning as the future path for strong, context-aware, and multilingual disambiguation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。