开源多模态模型可自动评估其他模型表现,媲美GPT且支持自我改进。
LLaVA-Critic: Learning to Evaluate Multimodal Models
- 用高质量评分数据训练,能理解多种评价标准和场景。
- 在多项评测中表现不逊于甚至超过GPT,可作为可靠评分器。
- 适合需要自动化评估与对齐训练的研究者使用。
我们提出 LLaVA-Critic,首个开源的大规模多模态模型(LMM),专为通用评估设计,可覆盖多种多模态任务。该模型基于高质量的评判指令跟随数据集进行训练,涵盖多样化的评价标准与场景。实验表明其在两个关键方向均具有效性:(1) 多模态模型自评(LMM-as-a-Judge),LLaVA-Critic 提供可靠评分,在多个评测基准上表现与或优于 GPT 模型;(2) 偏好学习(Preference Learning),生成用于偏好训练的奖励信号,提升模型对齐能力。本工作揭示了开源 LMM 在自我批判与评估中的潜力,为未来实现可扩展、超人类对齐反馈机制铺平道路。
原文摘要 · Abstract (English)
We introduce LLaVA-Critic, the first open-source large multimodal model (LMM) designed as a generalist evaluator to assess performance across a wide range of multimodal tasks. LLaVA-Critic is trained using a high-quality critic instruction-following dataset that incorporates diverse evaluation criteria and scenarios. Our experiments demonstrate the model's effectiveness in two key areas: (1) LMM-as-a-Judge, where LLaVA-Critic provides reliable evaluation scores, performing on par with or surpassing GPT models on multiple evaluation benchmarks; and (2) Preference Learning, where it generates reward signals for preference learning, enhancing model alignment capabilities. This work underscores the potential of open-source LMMs in self-critique and evaluation, setting the stage for future research into scalable, superhuman alignment feedback mechanisms for LMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。