arXiv:2506.07202cs.AI2025-06被引 2

通过动态任务扰动,揭示多模态模型是真懂还是只靠数据泄露答题。

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation

  • 用同一图像测试多种任务,看模型能否跨任务稳定表现。
  • 在MME、RealWorldQA等数据集上,模拟测试数据污染让模型性能飙升但泛化变差。
  • 适合关注模型真实理解力、避免被表面成绩误导的研究者使用。

多模态大语言模型在视觉-语言基准测试中表现优异,但训练过程中存在测试集暴露(数据污染)的风险,可能掩盖其真实泛化能力。这一问题在通过强化学习微调的推理型多模态模型中尤为突出。本文提出一种动态评估框架,超越传统静态基准,不扰动输入,而是扰动任务本身:对同一视觉输入,评估模型在问答、描述生成、提问、验证等多种任务上的表现,以探测其多样能力。该方法类比损失曲面的平坦性——在单一任务上过拟合或受污染的模型(尖锐极小值)在任务切换时表现崩溃,而具备泛化能力的模型(平坦极小值)则保持稳定。我们构建了自动化流水线,采用校准后的评判器,通过重述和破坏采样对开放式生成内容(如描述、问题)进行评分。应用于MME、RealWorldQA、CVRR-ES等基准上的主流图像/视频多模态模型,分析其跨任务的“能力向量”。结果表明,基于模拟测试数据的微调(极端污染)虽显著提升特定任务性能,却严重损害整体泛化能力。动态任务扰动为理解多模态模型的真实泛化能力提供了更深入洞察,可区分真正理解与虚假泄漏或过拟合。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) show impressive vision-language benchmark performance, yet growing concerns about data contamination (test set exposure during training) risk masking true generalization. This concern extends to reasoning MLLMs, often fine-tuned via reinforcement learning from potentially contaminated base models. We propose a novel dynamic evaluation framework to rigorously assess MLLM generalization, moving beyond static benchmarks. Instead of perturbing inputs, we perturb the task itself. Using the same visual input, models are evaluated across a family of tasks (e.g., QA, captioning, question posing, verification) to probe diverse capabilities. This task perturbation reveals whether model performance is robust or reliant on superficial task-specific cues. Our approach is analogous to loss landscape sharpness: models overfit or contaminated for a single task (sharp minima) falter under task shifts, unlike models with generalizable solutions (flatter minima). We developed an automated pipeline with a calibrated judge scoring open-ended generations (captions, questions) using paraphrase and corruption sampling. Applying this framework to leading image/video MLLMs on benchmarks including MME, RealWorldQA, and CVRR-ES, we analyze each model's cross-task "ability vector." We demonstrate that fine-tuning on simulated test data (extreme contamination) drastically sharpens task-specific performance but harms overall generalization. Our dynamic task perturbation offers deeper insights into MLLM generalization, distinguishing genuine understanding from spurious leakage or overfitting.

多模态模型评估泛化能力数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。