arXiv:2410.18071cs.CVcs.AI2024-10IJCAI被引 1

定制提示词让多模态大模型真实能力被挖掘

TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts

  • 按模型特性动态调整提示词,减少评估偏差
  • 实验证明可显著提升模型能力展现效果
  • 适合研究者构建更公平的多模态模型评测体系

近期,多模态大语言模型(MLLMs)因其出色能力受到广泛关注。对MLLMs的评估正变得愈发关键,有助于分析模型属性并提供有价值见解。然而,现有基准测试忽视了提示敏感性问题——微小的提示变化可能导致性能显著波动。不当提示可能掩盖模型真实能力,低估其表现。此外,不同模型对提示存在偏好差异,统一提示会导致评估偏差。本文分析了现有基准的这一缺陷,提出新的评估框架TP-Eval,引入提示定制化方法以降低偏差并挖掘模型潜力。TP-Eval将原始提示重写为针对不同模型定制的提示,并设计了专用于MLLM评估场景的模块。大量实验表明该方法能有效揭示模型真实能力,有助于社区构建更全面、可信的MLLM评估基准。

原文摘要 · Abstract (English)

Recently, multimodal large language models (MLLMs) have received much attention for their impressive capabilities. The evaluation of MLLMs is becoming critical to analyzing attributes of MLLMs and providing valuable insights. However, current benchmarks overlook the problem of prompt sensitivity - minor prompt variations may lead to significant performance fluctuations. Thus, inappropriate prompts may obscure the models' capabilities, underestimating the models' performance. Moreover, different models have different preferences for different prompts, and thus, using the same prompt for all models will cause evaluation bias. This paper analyzes this deficiency in existing benchmarks and further introduces a new evaluation framework named TP-Eval, which introduces a prompt customization method to reduce evaluation biases and tap models' potential. TP-Eval will rewrite the original prompts to different customized prompts for different models. In particular, we propose some well-designed modules for prompt customization tailored to the scenario of MLLM evaluation. Extensive experiments demonstrate the effectiveness of our approach to uncovering models' capabilities, and TP-Eval should benefit the community in developing more comprehensive and convincing MLLM evaluation benchmarks.

多模态提示工程模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。