评测多模态大模型在不同图像分辨率下的表现稳定性
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- 构建12个分辨率层级的综合评估基准,覆盖1.44万样本
- 提出相关性与误差指标,量化模型性能波动情况
- 适合关注模型鲁棒性与实际部署效果的研究者
多模态大语言模型(MLLMs)越来越多地支持动态图像分辨率输入。然而,当前评估范式主要关注语义性能,忽略了分辨率鲁棒性这一关键问题——模型性能是否在不同输入分辨率下保持稳定。为填补这一空白,我们提出 extbf{Res-Bench},一个包含14,400个样本、覆盖12个分辨率层级和六个核心能力维度的综合性基准。我们设计了一种新颖的评估框架,超越传统准确率指标,以捕捉性能稳定性。该框架引入多个鲁棒性度量:斯皮尔曼相关性用于评估分辨率-性能趋势,绝对/相对连续误差(ACE/RCE)用于衡量性能波动。利用这些指标,我们对主流MLLMs进行了大规模评估,分析包括:(1) 模型级与任务级鲁棒性考察,(2) 预处理策略(如填充与超分辨率)的影响,(3) 微调对稳定性提升的探索。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robustness - whether performance remains stable across varying input resolutions. To address this gap, we introduce \textbf{Res-Bench}, a comprehensive benchmark comprising 14,400 samples across 12 resolution levels and six core capability dimensions. We designed a novel evaluation framework that goes beyond traditional accuracy metrics to capture performance stability. This framework introduces multiple robustness metrics: Spearman's correlation for assessing resolution-performance trends, and Absolute/Relative Continuous Error (ACE/RCE) for measuring performance volatility. Using these metrics, we conducted a large-scale evaluation of leading MLLMs. Our analysis encompasses: (1) model-centric and task-centric robustness examination, (2) investigation of preprocessing strategies including padding and super-resolution, and (3) exploration of fine-tuning for stability enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。