提出可解释的多模态评估模型,实现跨任务可靠打分。
Judge Model for Large-scale Multimodality Benchmarks
- 构建多模态判断模型,聚合跨模态意见并分析推理一致性。
- 在280个样本上与人类评分高度一致,相关性达0.92以上。
- 适合大规模多模态模型评测,尤其适用于需要可解释性的研究。
我们提出一种专用的多模态判断模型,用于在多样化任务中提供可靠且可解释的评估。该基准涵盖文本、音频、图像和视频模态,基于精心采样的公开数据集,并采用固定种子以确保可复现性并减少训练-测试泄露。不同于简单打分,该框架聚合多模态判断,分析模型输出的质量与推理一致性,并生成诊断反馈。我们在280个多模态样本上评估了多个多模态大模型(包括Gemini 2.5、Phi 4和Qwen 2.5),并将判断模型评分与人工标注结果对比。结果显示,判断模型与人类评分高度一致,证明其在未来的多模态人工智能研究中具备作为可扩展、可解释评估流程的潜力。
原文摘要 · Abstract (English)
We propose a dedicated multimodal Judge Model designed to provide reliable, explainable evaluation across a diverse suite of tasks. Our benchmark spans text, audio, image, and video modalities, drawing from carefully sampled public datasets with fixed seeds to ensure reproducibility and minimize train test leakage. Instead of simple scoring, our framework aggregates multimodal judgments, analyzes the quality and reasoning consistency of model outputs, and generates diagnostic feedback. We evaluate several MLLMs, including Gemini 2.5, Phi 4, and Qwen 2.5, across 280 multimodal samples and compare judge model assessments with human annotators. Results show strong alignment between the Judge Model and human scores, demonstrating its potential as a scalable, interpretable evaluation pipeline for future multimodal AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。