用视觉蕴含任务检验多模态模型理解能力,发现其既有效又存陷阱。
Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
- 通过零样本、少样本和微调对比,测试提示设计对模型表现的影响。
- 微调后在e-SNLI-VE上达83.3%准确率,优于当前最优OFA-X模型。
- 模型常凭语言先验编造视觉内容,解释结果虽似人但难证真视觉理解。
本研究以LLaMA 3.2 11B Vision模型为案例,探究视觉蕴含(VE)任务作为多模态语言模型视觉-语言理解能力诊断工具的可靠性。实验涵盖零样本、少样本及微调设置,考察提示设计、上下文示例数量与顺序、视觉信息可用性等因素对性能的影响。结果表明,三样本推理优于零样本基线,但增加示例会引入更多噪声。标签顺序显著影响预测结果。缺乏视觉信息时,模型表现出强烈幻觉倾向,依赖语言先验。微调后在e-SNLI-VE数据集上达到83.3%准确率,超越现有最优模型OFA-X。解释评估显示,微调模型生成语义合理的解释,BERTScore F1达89.2%。然而,在视觉信息受限条件下也获得相近的BERTScore,质疑该任务的视觉根基性。整体表明VE任务兼具诊断价值与局限,需改进多模态评估方法。
原文摘要 · Abstract (English)
This study investigates the extent to which the Visual Entailment (VE) task serves as a reliable probe of vision-language understanding in multimodal language models, using the LLaMA 3.2 11B Vision model as a test case. Beyond reporting performance metrics, we aim to interpret what these results reveal about the underlying possibilities and limitations of the VE task. We conduct a series of experiments across zero-shot, few-shot, and fine-tuning settings, exploring how factors such as prompt design, the number and order of in-context examples and access to visual information might affect VE performance. To further probe the reasoning processes of the model, we used explanation-based evaluations. Results indicate that three-shot inference outperforms the zero-shot baselines. However, additional examples introduce more noise than they provide benefits. Additionally, the order of the labels in the prompt is a critical factor that influences the predictions. In the absence of visual information, the model has a strong tendency to hallucinate and imagine content, raising questions about the model's over-reliance on linguistic priors. Fine-tuning yields strong results, achieving an accuracy of 83.3% on the e-SNLI-VE dataset and outperforming the state-of-the-art OFA-X model. Additionally, the explanation evaluation demonstrates that the fine-tuned model provides semantically meaningful explanations similar to those of humans, with a BERTScore F1-score of 89.2%. We do, however, find comparable BERTScore results in experiments with limited vision, questioning the visual grounding of this task. Overall, our results highlight both the utility and limitations of VE as a diagnostic task for vision-language understanding and point to directions for refining multimodal evaluation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。