无需标签即可对大型多模态模型进行性能排序。
Ranked from Within: Ranking Large Multimodal Models Without Labels
- 利用模型输出的不确定性得分,实现无监督模型排序。
- 47个主流多模态模型在9个视觉问答任务中表现一致。
- 适合在无标注数据上快速选型的场景,节省人工标注成本。
预训练大型多模态模型(LMM)的性能能否在无标签情况下预测?随着LMM数量激增,面对新数据或任务时高效选择模型变得愈发重要。传统方法相当于给模型考试并打分,需人工标注答案。本文摒弃此法,转而分析模型自身输出的信号,考察其对自身能力边界的认知程度,评估这些信号在无监督模型排序中的有效性。我们在9个视觉问答基准上评测了47个最先进的LMM(如LLaVA),发现基于softmax分布的不确定性分数能稳健且一致地预测模型相对性能。该方法可在无标签数据上实现LMM排序,为多样化目标领域提供无需人工标注的模型选择方案。
原文摘要 · Abstract (English)
Can the relative performance of a pre-trained large multimodal model (LMM) be predicted without access to labels? As LMMs proliferate, it becomes increasingly important to develop efficient ways to choose between them when faced with new data or tasks. The usual approach does the equivalent of giving the models an exam and marking them. We opt to avoid marking and the associated labor of determining the ground-truth answers. Instead, we explore other signals elicited and ascertain how well the models know their own limits, evaluating the effectiveness of these signals at unsupervised model ranking. We evaluate $47$ state-of-the-art LMMs (\eg, LLaVA) across $9$ visual question answering benchmarks, analyzing how well uncertainty-based metrics can predict relative model performance. Our findings show that uncertainty scores derived from softmax distributions provide a robust and consistent basis for ranking models across various tasks. This facilitates the ranking of LMMs on unlabeled data, providing a practical approach for selecting models for diverse target domains without requiring manual annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。