arXiv:2503.05977cs.CVcs.AI2025-03ICLR被引 16

用多个视频语言模型评估模型,结果反而更不可靠。

Is Your Video Language Model a Reliable Judge?

  • 让多个VLM共同打分,试图提升评估可靠性。
  • 混合可靠与不可靠模型打分,整体准确率不升反降。
  • 单纯提升理解力不够,需更智能的可信度评估机制。

随着视频语言模型(VLMs)在各类场景中应用增多,其性能评估的稳健性与可扩展性日益关键。传统人工专家评估存在一致性差、规模小的问题,促使研究转向使用VLM自动评估自身。然而,现有方法多依赖单一VLM作为评判者,易受理解能力不足或固有偏见影响,导致评估不可靠。本研究探索集体判断策略:通过聚合多个VLM的评分来提升可靠性。结果发现,当评估池中包含不可靠模型时,引入它们会带来噪声,反而降低最终评估的准确性。进一步对表现较差的Video-LLaVA进行微调,发现仅提升理解能力仍不足以使其成为可靠评判者。该研究揭示了集体思维方法的局限性,强调亟需能识别并加权个体模型可信度的新评估范式。

原文摘要 · Abstract (English)

As video language models (VLMs) gain more applications in various scenarios, the need for robust and scalable evaluation of their performance becomes increasingly critical. The traditional human expert-based evaluation of VLMs has limitations in consistency and scalability, which sparked interest in automatic methods such as employing VLMs to evaluate VLMs. However, the reliability of VLMs as judges remains underexplored. Existing methods often rely on a single VLM as the evaluator. However, this approach can be unreliable or biased because such a model may lack the ability to fully understand the content and may have inherent biases, ultimately compromising evaluation reliability. A remedy is to apply the principle of collective thoughts, aggregating evaluations from multiple VLMs to enhance reliability. This study investigates the efficacy of such approaches, particularly when the pool of judges includes both reliable and unreliable models. Our findings reveal that incorporating collective judgments from such a mixed pool does not necessarily improve the accuracy of the final evaluation. The inclusion of less reliable judges can introduce noise, undermining the overall reliability of the outcomes. To explore the factors that impact evaluation reliability, we fine-tune an underperforming VLM judge, Video-LLaVA, and observe that improved understanding ability alone is insufficient to make VLM judges more reliable. These findings stress the limitations of collective thought approaches and highlight the need for more advanced methods that can account for the reliability of individual models. Our study promotes the development of more reliable evaluation methods for VLMs

视频评估模型评测多模型决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。