用大模型生成的错觉事实,识别图像是否违背常识。
Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts
- 通过大模型提取图像原子事实,利用自然语言推理判断真伪
- 在WHOOPS!数据集零样本下达到新最好性能
- 适合需要检测图像常识性错误的研究与应用
衡量图像真实性仍是人工智能领域的难题。例如,爱因斯坦手持智能手机违反常识,因现代智能手机在他去世后才出现。本文提出一种基于大视觉语言模型(LVLM)和自然语言推理(NLI)的新方法评估图像真实性。该方法基于假设:当面对违背常识的图像时,LVLM会产生幻觉。通过LVLM从图像中提取原子事实,得到真实与错误事实的混合结果。计算这些事实之间的成对蕴含分数,并聚合为单一现实性得分,以识别真实事实与幻觉之间的矛盾,从而揭示违背常识的图像。该方法在WHOOPS!数据集的零样本设置下达到了新的最先进性能。
原文摘要 · Abstract (English)
Quantifying the realism of images remains a challenging problem in the field of artificial intelligence. For example, an image of Albert Einstein holding a smartphone violates common-sense because modern smartphone were invented after Einstein's death. We introduce a novel method for assessing image realism using Large Vision-Language Models (LVLMs) and Natural Language Inference (NLI). Our approach is based on the premise that LVLMs may generate hallucinations when confronted with images that defy common sense. Using LVLM to extract atomic facts from these images, we obtain a mix of accurate facts and erroneous hallucinations. We proceed by calculating pairwise entailment scores among these facts, subsequently aggregating these values to yield a singular reality score. This process serves to identify contradictions between genuine facts and hallucinatory elements, signaling the presence of images that violate common sense. Our approach has achieved a new state-of-the-art performance in zero-shot mode on the WHOOPS! dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。