用可解释的视觉概念生成更准确的阿拉伯语图像描述。
Multimodal Arabic Captioning with Interpretable Visual Concept Integration
- 通过多语言编码器提取阿拉伯语可解释视觉概念
- mCLIP+Gemini组合在BLEU-1达5.34%,相似度60.01%
- 适合需要文化语境准确性的阿拉伯语图文生成场景
我们提出VLCAP框架,结合CLIP-based视觉标签检索与多模态文本生成实现阿拉伯语图像描述。不依赖端到端方法,而是通过mCLIP、AraCLIP和Jina V4三个多语言编码器分别提取可解释的阿拉伯语视觉概念。构建包含约21,000个通用领域标签的混合词汇表,这些标签源自Visual Genome数据集的翻译。将前k个检索到的标签转化为流畅的阿拉伯语提示,并与原图一同输入视觉语言模型。第二阶段测试了Qwen-VL和Gemini Pro Vision,形成六种编码器-解码器组合。结果显示,mCLIP + Gemini Pro Vision在BLEU-1上达到5.34%,余弦相似度为60.01%;AraCLIP + Qwen-VL获得最高LLM评分36.33%。该可解释流程生成符合文化背景且上下文准确的阿拉伯语描述。
原文摘要 · Abstract (English)
We present VLCAP, an Arabic image captioning framework that integrates CLIP-based visual label retrieval with multimodal text generation. Rather than relying solely on end-to-end captioning, VLCAP grounds generation in interpretable Arabic visual concepts extracted with three multilingual encoders, mCLIP, AraCLIP, and Jina V4, each evaluated separately for label retrieval. A hybrid vocabulary is built from training captions and enriched with about 21K general domain labels translated from the Visual Genome dataset, covering objects, attributes, and scenes. The top-k retrieved labels are transformed into fluent Arabic prompts and passed along with the original image to vision-language models. In the second stage, we tested Qwen-VL and Gemini Pro Vision for caption generation, resulting in six encoder-decoder configurations. The results show that mCLIP + Gemini Pro Vision achieved the best BLEU-1 (5.34%) and cosine similarity (60.01%), while AraCLIP + Qwen-VL obtained the highest LLM-judge score (36.33%). This interpretable pipeline enables culturally coherent and contextually accurate Arabic captions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。