用SelfCheckGPT检测大模型交通图像描述中的幻觉,提升准确性。
LLMs Can Check Their Own Results to Mitigate Hallucinations in Traffic Understanding Tasks
- 引入SelfCheckGPT机制,让大模型自检生成内容真伪。
- 白天图像检测效果优于黎明、黄昏和夜间,准确率更高。
- GPT-4o生成描述更真实,但误判非幻觉内容为幻觉较多。
当前大型语言模型(LLMs)在文本生成到多模态处理方面表现出色,正被探索用于车载系统中的感知任务,如高级驾驶辅助系统(ADAS)或自动驾驶(AD)。然而,这些模型常产生不连贯或不真实的“幻觉”信息。本文系统评估了SelfCheckGPT在三款先进LLM(GPT-4o、LLaVA、Llama3)分析来自美国Waymo Open Dataset与瑞典PREPER CITY数据集的视觉汽车数据时,识别幻觉的能力。结果显示,GPT-4o生成图像描述更忠实于原图,而其对非幻觉内容误判为幻觉的比例高于LLaVA。不同数据集类型(Waymo或PREPER CITY)对描述质量或幻觉检测效果无显著影响,但模型在白天图像上的表现优于黎明、黄昏和夜间。总体表明,SelfCheckGPT及其适配方法可用于过滤顶尖大模型在交通图像描述中产生的幻觉。
原文摘要 · Abstract (English)
Today's Large Language Models (LLMs) have showcased exemplary capabilities, ranging from simple text generation to advanced image processing. Such models are currently being explored for in-vehicle services such as supporting perception tasks in Advanced Driver Assistance Systems (ADAS) or Autonomous Driving (AD) systems, given the LLMs' capabilities to process multi-modal data. However, LLMs often generate nonsensical or unfaithful information, known as ``hallucinations'': a notable issue that needs to be mitigated. In this paper, we systematically explore the adoption of SelfCheckGPT to spot hallucinations by three state-of-the-art LLMs (GPT-4o, LLaVA, and Llama3) when analysing visual automotive data from two sources: Waymo Open Dataset, from the US, and PREPER CITY dataset, from Sweden. Our results show that GPT-4o is better at generating faithful image captions than LLaVA, whereas the former demonstrated leniency in mislabeling non-hallucinated content as hallucinations compared to the latter. Furthermore, the analysis of the performance metrics revealed that the dataset type (Waymo or PREPER CITY) did not significantly affect the quality of the captions or the effectiveness of hallucination detection. However, the models showed better performance rates over images captured during daytime, compared to during dawn, dusk or night. Overall, the results show that SelfCheckGPT and its adaptation can be used to filter hallucinations in generated traffic-related image captions for state-of-the-art LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。