用程序化方法评估视觉语言模型回答的真实性与有用性
Trust but Verify: Programmatic VLM Evaluation in the Wild
- 用场景图+程序验证构建可自动检验的问答数据集
- 10.5k问答对中多数模型难以兼顾回答准确与有用
- 适合关注模型幻觉、评测可靠性的研究者使用
视觉语言模型(VLM)常对视觉问题生成看似合理但错误的回答。然而,在开放式问答中可靠量化此类幻觉的影响极具挑战,因需对每个回答进行视觉验证。本文提出程序化VLM评估(PROVE),一种针对开放式查询的新型评测范式。通过将高保真场景图(由超详细图像描述构建)提供给大语言模型(LLM),并引导其生成多样化的问答对及可在场景图上执行的验证程序,构建了包含10.5k个具有挑战性且视觉可验证的问答对的基准。进一步提出基于场景图的统一框架,程序化评估模型自由回答在有用性和真实性上的表现。在该基准上测试多个VLM,发现极少模型能同时实现良好平衡。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) often generate plausible but incorrect responses to visual queries. However, reliably quantifying the effect of such hallucinations in free-form responses to open-ended queries is challenging as it requires visually verifying each claim within the response. We propose Programmatic VLM Evaluation (PROVE), a new benchmarking paradigm for evaluating VLM responses to open-ended queries. To construct PROVE, we provide a large language model (LLM) with a high-fidelity scene-graph representation constructed from a hyper-detailed image caption, and prompt it to generate diverse question-answer (QA) pairs, as well as programs that can be executed over the scene graph object to verify each QA pair. We thus construct a benchmark of 10.5k challenging but visually grounded QA pairs. Next, to evaluate free-form model responses to queries in PROVE, we propose a programmatic evaluation strategy that measures both the helpfulness and truthfulness of a response within a unified scene graph-based framework. We benchmark the helpfulness-truthfulness trade-offs of a range of VLMs on PROVE, finding that very few are in-fact able to achieve a good balance between the two. Project page: \url{https://prove-explorer.netlify.app/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。