用结构方程模型构建更科学的多模态大模型评估基准
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
- 基于皮亚杰能力层级,将模型能力分为感知、记忆、推理三类
- 新基准GOLD显著降低指标冗余,提升认知一致性与可解释性
- 适合研究模型内在能力结构与评估体系的学者参考
评估多模态大语言模型(MLLMs)面临根本性挑战:缺乏结构化、可解释且有理论基础的评测基准;现有任务多为启发式分组,存在认知目标模糊、能力重叠、指标冗余和诊断力弱等问题。为此,我们提出一种基于结构方程模型的对齐框架,量化内部效度、维度分离性和组件贡献度,并引入受皮亚杰启发的能力层级,将MLLM能力划分为感知、记忆和推理三个层次。在此理论基础上重新组织现有任务,构建GOLD基准。实验表明,该基准在可解释性、指标冗余程度和认知一致性方面均优于以往基准。
原文摘要 · Abstract (English)
Evaluating multimodal large language models (MLLMs) is fundamentally challenged by the absence of structured, interpretable, and theoretically grounded benchmarks; current heuristically-grouped tasks have vague cognitive targets, overlapping abilities, redundant indicators, and weak diagnostic power. We therefore propose a structural-equation-modeling-aligned framework that quantifies internal validity, dimensional separability, and component contributions, and introduce a Piaget-inspired capability hierarchy that stratifies MLLM abilities into Perception, Memory, and Reasoning. Reorganizing existing tasks under this theory, we build the GOLD benchmark, whose experiments show superior interpretability, lower indicator redundancy, and clearer cognitive consistency than prior benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。