无需人工标注,让多模态模型互评答题质量。
UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation
- 用图像自动生成问题,模型间互相评审答案。
- 在MMstar上与人类评估相关性达0.944,接近人工水平。
- 适合想高效评估多模态大模型的科研人员使用。
多模态大语言模型(MLLMs)在视觉问答(VQA)任务中崭露头角,推动了客观评估方法的研究。现有评估方法受限于人工设计问答对的工作量,难以大规模开展;而自动化“模型作为评判者”的方法常引入偏差。为此,我们提出无监督同伴评审框架UPME,仅依赖图像数据,使模型能自动生成问题并由其他模型互评答案,显著减少对人工干预的依赖。同时,引入视觉-语言评分体系,从三个方面评估:(i) 回答正确性;(ii) 视觉理解与推理能力;(iii) 图像-文本关联性。实验表明,UPME在MMstar数据集上与人工评估的皮尔逊相关系数达0.944,在ScienceQA数据集上为0.814,证明该框架与人工基准及人类偏好高度一致。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have emerged to tackle the challenges of Visual Question Answering (VQA), sparking a new research focus on conducting objective evaluations of these models. Existing evaluation methods face limitations due to the significant human workload required to design Q&A pairs for visual images, which inherently restricts the scale and scope of evaluations. Although automated MLLM-as-judge approaches attempt to reduce the human workload through automatic evaluations, they often introduce biases. To address these problems, we propose an Unsupervised Peer review MLLM Evaluation framework. It utilizes only image data, allowing models to automatically generate questions and conduct peer review assessments of answers from other models, effectively alleviating the reliance on human workload. Additionally, we introduce the vision-language scoring system to mitigate the bias issues, which focuses on three aspects: (i) response correctness; (ii) visual understanding and reasoning; and (iii) image-text correlation. Experimental results demonstrate that UPME achieves a Pearson correlation of 0.944 with human evaluations on the MMstar dataset and 0.814 on the ScienceQA dataset, indicating that our framework closely aligns with human-designed benchmarks and inherent human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。