arXiv:2506.01738cs.CV2025-06

构建首个通用视觉评分基准STORM,评估多模态大模型在有序回归任务中的表现。

STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset

  • 提出粗到细处理流程,动态考虑评分候选并生成可解释思维链。
  • 涵盖14个数据集、65.5万张图像对,覆盖5类视觉评分场景。
  • 适合研究模型在零样本场景下理解有序评分关系的能力。

视觉评分是人工智能对视觉内容进行多维量化的重要能力,广泛应用于图像质量评估、人脸年龄估计和医学影像分级等有序回归任务。然而当前多模态大模型在该能力上表现不佳,且缺乏相关数据集与评测基准。本文构建并发布STORM——一个面向多模态大模型视觉评分能力的综合性基准。STORM包含跨5个常见视觉评分领域的14个有序回归数据集,共65.5万张图像级配对及其精心标注的视觉问答(VQA)。我们还提出一种粗到细处理流程,动态考虑标签候选并生成可解释的推理过程,为模型提供通用且可信的有序思考范式。该基准旨在评估多模态大模型在需要理解评分标签本质有序关系场景下的全栈式零样本性能。大量实验验证了框架有效性,并为更优微调策略提供启示。数据集、代码与预训练模型已开源:https://storm-bench.github.io/。

原文摘要 · Abstract (English)

Visual rating is an essential capability of artificial intelligence (AI) for multi-dimensional quantification of visual content, primarily applied in ordinal regression (OR) tasks such as image quality assessment, facial age estimation, and medical image grading. However, current multi-modal large language models (MLLMs) under-perform in such visual rating ability while also suffering the lack of relevant datasets and benchmarks. In this work, we collect and present STORM, a data collection and benchmark for Stimulating Trustworthy Ordinal Regression Ability of MLLMs for universal visual rating. STORM encompasses 14 ordinal regression datasets across five common visual rating domains, comprising 655K image-level pairs and the corresponding carefully curated VQAs. Importantly, we also propose a coarse-to-fine processing pipeline that dynamically considers label candidates and provides interpretable thoughts, providing MLLMs with a general and trustworthy ordinal thinking paradigm. This benchmark aims to evaluate the all-in-one and zero-shot performance of MLLMs in scenarios requiring understanding of the essential common ordinal relationships of rating labels. Extensive experiments demonstrate the effectiveness of our framework and shed light on better fine-tuning strategies. The STORM dataset, benchmark, and pre-trained models are available on the following webpage to support further research in this area. Datasets and codes are released on the project page: https://storm-bench.github.io/.

视觉评分多模态有序回归基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。