提出无需标注的视觉语言逻辑一致性评估方法,解决大模型验证难题。
Towards Annotation-Free Validation of MLLMs: A Vision-Language Logical Consistency Metric

- 基于因果逻辑构建视觉语言一致性度量,无需真实答案即可评估。
- 11个主流多模态模型在逻辑一致性上显著落后于准确率表现。
- 适合用于新任务下的模型筛选与可靠回答验证,尤其无标注场景。
现有主流评估方法可能奖励大语言模型的盲目猜测,且在缺乏真实标注时难以适用于新任务的模型验证。基于基本逻辑原则,我们提出一种新框架,评估多模态大模型在充分与必要因果关系上的视觉-语言逻辑一致性。定义了视觉-语言逻辑一致性度量(VL-LCM),应用于传统MC-VQA测试和无需真实标注的NaturalBench测试。在代表性视觉语言基准MMMU及Recent NaturalBench挑战上,对4大前沿家族的11个开源多模态大模型进行系统评估。结果表明,尽管近期模型在准确率上取得显著进步,其逻辑一致性仍明显滞后。通过分析VL-LCM与带标注指标的相关性、度量可靠性及与回答分布的关系,验证了其在无标注情况下的有效性与适用性。研究建议:除准确率外,逻辑一致性可同时用于评估准确性与可靠性,且可用于无标注新任务中的模型选择、验证与可信回答解释。
原文摘要 · Abstract (English)
Dominant accuracy evaluation might reward unwarranted guessing of Large Language Models, and it might not be applicable to novel tasks for model validation without ground-truth (gt) annotation. Based on basic logic principle, we propose a novel framework to evaluate the vision-language logical consistency of MLLMs on both sufficient and necessary cause-effect relations. We define Vision-Language Logical Consistency Metric (VL-LCM) on traditional MC-VQA tests, and recent NaturalBench tests without the need for gt annotation. Through systematic experiments on representative VL benchmark MMMU and recent VL challenges like NaturalBench, we evaluated 11 recent open-source MLLMs from 4 frontier families. Our findings reveal that, despite significant progress of recent MLLMs on accuracy, logical consistency lags behind significantly. Extensive evaluations on the correlations of VL-LCM with metrics on gt, the reliability of LCM, and the relation of VL-LCM with response distribution justify the validity and applicability of VL-LCM even without gt annotation. Our findings suggest that, beyond accuracy, logical consistency could be employed for both accuracy and reliability. VL-LCM can also be employed for MLLM selection, validation, and reliable answer justification in novel tasks without gt annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。