arXiv:2505.22305cs.CV2025-05中稿 · DIS'25被引 2

用可视化热图评估视觉语言模型可靠性,无需真实标签

IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth

  • 将模型输出转为绿红热图,直观显示物体存在与否
  • 引入'间谍物体'检测模型幻觉,发现模型误判
  • 用户仅看少量热图单元就能做出可靠判断

我们提出IKIWISI(“我一看就知道”),一种用于视频目标识别中无真实标签时评估视觉语言模型可靠性的交互式视觉模式生成工具。IKIWISI将模型输出转化为二值热图,绿色表示物体存在,红色表示不存在。该可视化利用人类天生的模式识别能力来评估模型可靠性。IKIWISI引入“间谍物体”——用户明确知道不存在的对抗性实例,以识别模型对不存在项目的幻觉。该工具作为认知审计机制,通过可视化模型与人类理解之间的偏差,揭示两者不一致之处。15名参与者的研究表明,用户认为IKIWISI易用,其评估结果在有客观指标时与之相关,并能仅通过观察少量热图单元即得出合理结论。该方法不仅通过自定义物体集的视觉评估补充了传统评价手段,还揭示了提升人类感知与机器理解对齐度的潜在路径。

原文摘要 · Abstract (English)

We present IKIWISI ("I Know It When I See It"), an interactive visual pattern generator for assessing vision-language models in video object recognition when ground truth is unavailable. IKIWISI transforms model outputs into a binary heatmap where green cells indicate object presence and red cells indicate object absence. This visualization leverages humans' innate pattern recognition abilities to evaluate model reliability. IKIWISI introduces "spy objects": adversarial instances users know are absent, to discern models hallucinating on nonexistent items. The tool functions as a cognitive audit mechanism, surfacing mismatches between human and machine perception by visualizing where models diverge from human understanding. Our study with 15 participants found that users considered IKIWISI easy to use, made assessments that correlated with objective metrics when available, and reached informed conclusions by examining only a small fraction of heatmap cells. This approach not only complements traditional evaluation methods through visual assessment of model behavior with custom object sets, but also reveals opportunities for improving alignment between human perception and machine understanding in vision-language systems.

视觉语言模型模型评估可视化幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。