构建首个全面评估多模态大模型上下文适应能力的基准
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

- 设计涵盖41000+图像-问题对的多场景测试集
- 发现现有模型在上下文偏移时难以平衡回答与拒绝
- 适合研究多模态模型鲁棒性与安全性的人群使用
多模态大语言模型在视觉-语言任务中表现优异,但在上下文不完整或发生偏移时常失效。可靠的模型应能拒绝真正上下文无关的问题(主题级偏移),同时仍可回答非主题偏移的上下文偏移问题。现有基准主要关注上下文无关或视觉无法解答的问题,忽视了可回答的上下文偏移情况,且覆盖的偏移类型有限。为此,我们提出MMOOC,一个大规模基准,用于评估多模态大模型的拒绝能力和鲁棒回答能力。MMOOC包含超过41,000个图像-问题对,涵盖三种问题形式、八种偏移类型和六种视觉场景,通过多模态大模型过滤与人工验证确保数据质量。我们采用准确率与拒绝率评估模型表现,并引入大模型作为裁判的评估指标以判断推理正确性。对多种多模态大模型的实验表明,当前模型在上下文偏移下仍难以平衡回答与拒绝能力。我们进一步分析关键失败模式,并发现后训练可提升模型鲁棒性。MMOOC将公开发布。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。