测试大模型在真实场景中对虚假关联的泛化能力,发现主流模型表现不佳。
Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
- 基于真实VQA错误构建含124类虚假关联的新基准
- 最先进闭源模型在新基准上准确率仅37.1%
- 用合成数据微调可提升至78.4%,说明能学着避开捷径
微调可能引发非关键特征与目标标签间的虚假关联,但现有评测设置多为人为构造且任务范围狭窄。本文关注在大规模多模态视觉语言模型(LVLM)中,于无显式任务监督的海量多样化数据上预训练后产生的虚假关联。我们通过收集GPT-4o在真实世界视觉问答(VQA)基准上的错误,结合LVLM与人类标注及合成反事实评估,筛选出由虚假关联导致的错误,构建了名为SpuriVerse的新基准。该基准包含124种不同的虚假关联类型,每类含1个真实样本和10个合成样本,共1364道多选题。我们在SpuriVerse上评估了15个开源与闭源的LVLM,发现即使最先进的闭源模型表现仍显著不足,最高准确率仅为37.1%。对强调虚假关联的合成样本进行微调后,性能提升至78.40%,表明模型能从多样化的虚假模式中学习,在未见情境下避免使用捷径,转而关注整体图像上下文。
原文摘要 · Abstract (English)
Finetuning can cause spurious correlations to arise between non-essential features and the target labels, but benchmarks to study these effects involve contrived settings and narrow tasks. In contrast, we consider spurious correlations in multi-modal Large Vision Language Models (LVLMs) pretrained on extensive and diverse datasets without explicit task supervision. We develop a benchmark by sourcing GPT-4o errors on real-world visual-question-answering (VQA) benchmarks, then curating a subset through LVLM-human annotation and synthetic counterfactual evaluation to identify errors caused by spurious correlations. This process yields SpuriVerse, a novel benchmark comprised of 124 distinct types of spurious correlations extracted from real-world datasets, each containing 1 realistic and 10 synthetic VQA samples for a total of 1364 multiple choice questions. We evaluate 15 open and closed-source LVLMs on SpuriVerse, finding that even state-of-the-art closed-source models struggle significantly, achieving at best only 37.1% accuracy. Fine-tuning on synthetic examples that emphasize the spurious correlation improves performance to 78.40%, suggesting that training on diverse spurious patterns generalizes to unseen situations: models appear to learn to avoid "shortcuts" and attend to the overall image context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。