用新方法筛选出真正需要多模态推理的题目,让模型评估更准确。
Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory
- 将模型能力与题目难度拆分为图像、文本和跨模态三部分
- 在三个数据集上验证,即使一半题目是低质题也能保持排名可靠
- 适合需要高质量多模态评估的研究者使用
多模态大语言模型(MLLMs)近年来成为能处理多种模态的通用架构。评估其跨模态整合能力的基准应具备可靠性。然而,现有基准充斥着可通过单一模态解答的捷径题目,导致评估结果不可靠。例如,在视觉-语言任务中,仅凭图像或文本即可得出正确答案。这类低质量题目无谓增加基准规模与计算开销。本文提出多模态多维度项目反应理论框架(M3IRT),通过将模型能力与题目难度分解为图像独有、文本独有及跨模态三部分,实现对多模态推理能力的精准估计。该方法可识别并优先保留真正的跨模态题目,构建紧凑且高质量的评估子集。在24个视觉语言模型(VLMs)上测试显示,即便50%题目为人工生成的低质题,其排名保真度依然稳定,显著降低评估成本同时提升可靠性。M3IRT为评估跨模态推理与优化多模态基准提供了实用工具。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their ability for cross-modal integration. However, current benchmarks are filled with shortcut questions, which can be solved using only a single modality, thereby yielding unreliable rankings. For example, in vision-language cases, we can find the correct answer without either the image or the text. These low-quality questions unnecessarily increase the size and computational requirements of benchmarks. We introduce a multi-modal and multidimensional item response theory framework (M3IRT) that extends classical IRT by decomposing both model ability and item difficulty into image-only, text-only, and cross-modal components. M3IRT estimates cross-modal ability of MLLMs and each question's cross-modal difficulty, enabling compact, high-quality subsets that better reflect multimodal reasoning. Across 24 VLMs on three benchmarks, M3IRT prioritizes genuinely cross-modal questions over shortcuts and preserves ranking fidelity even when 50% of items are artificially generated low-quality questions, thereby reducing evaluation cost while improving reliability. M3IRT thus offers a practical tool for assessing cross-modal reasoning and refining multimodal benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。