arXiv:2604.01181cs.HCcs.CL2026-04

测试大模型识别数据可视化谎言的能力与意图判断

True (VIS) Lies: Analyzing How Generative AI Recognizes Intentionality, Rhetoric, and Misleadingness in Visualization Lies

  • 用修辞理论和作者意图分类法分析误导性图表
  • 16个模型在2336条新冠推文图中表现参差
  • 对比专家与模型判断,揭示人机差异

本研究探究多模态大语言模型识别和解释误导性可视化的能力,以及对其背后成因和潜在意图的辨识。分析基于可视化修辞理论与新构建的作者意图分类体系。通过包含2336条新冠相关推文(半数含误导性图表)的数据集,并补充来自VisLies(IEEE VIS活动)的真实案例,涵盖感知、认知与概念性错误。评估了16个前沿模型,包括小模型(如Nemotron-Nano-V2-VL,12B参数)、中型模型(如Qwen3-VL,235B参数)及大型模型(如Kimi-K2.5,1000B参数),并包含前沿闭源模型GPT-5.4。同时开展专家用户研究,比较模型与人类对同一误导性图表的修辞技巧与意图判断,揭示二者在理解上的异同,为理解模型认知边界提供依据。

原文摘要 · Abstract (English)

This study investigates the ability of multimodal Large Language Models (LLMs) to identify and interpret misleading visualizations, and recognize these observations along with their underlying causes and potential intentionality. Our analysis leverages concepts from visualization rhetoric and a newly developed taxonomy of authorial intents as explanatory lenses. We formulated three research questions and addressed them experimentally using a dataset of 2,336 COVID-19-related tweets, half of which contain misleading visualizations, and supplemented it with real-world examples of perceptual, cognitive, and conceptual errors drawn from VisLies, the IEEE VIS community event dedicated to showcasing deceptive and misleading visualizations. To ensure broad coverage of the current LLM landscape, we evaluated 16 state-of-the-art models. Among them, 15 are open-weight models, spanning a wide range of model sizes, architectural families, and reasoning capabilities. The selection comprises small models, namely Nemotron-Nano-V2-VL (12B parameters), Mistral-Small-3.2 (24B), DeepSeek-VL2 (27B), Gemma3 (27B), and GTA1 (32B); medium-sized models, namely Qianfan-VL (70B), Molmo (72B), GLM-4.5V (108B), LLaVA-NeXT (110B), and Pixtral-Large (124B); and large models, namely Qwen3-VL (235B), InternVL3.5 (241B), Step3 (321B), Llama-4-Maverick (400B), and Kimi-K2.5 (1000B). In addition, we employed OpenAI GPT-5.4, a frontier proprietary model. To establish a human perspective on these tasks, we also conducted a user study with visualization experts to assess how people perceive rhetorical techniques and the authorial intentions behind the same misleading visualizations. This allows comparison between model and expert behavior, revealing similarities and differences that provide insights into where LLMs align with human judgment and where they diverge.

可视化大模型意图识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。