arXiv:2604.21134cs.CL2026-04

让可视化智能体读懂图表数据,解决误读和混淆问题。

Beyond Pixels: Introspective and Interactive Grounding for Visualization Agents

论文配图:Beyond Pixels: Introspective and Interactive Grounding for Visualization Agents
图 1 · 摘自论文原文
  • 通过查询图表底层数据规范获取确定性证据,突破仅靠像素理解的局限。
  • 结合交互操作消除视觉模糊,在重叠图形上提升6.7%问答准确率。
  • 适用于需要精准分析图表的自动化数据探索与人机协同场景。

视觉语言模型(VLMs)常错误解读图表数值、虚构细节或混淆重叠元素。现有方法仅依赖像素解析,形成‘仅像素瓶颈’:将动态交互图表当作静态图像处理,丢失了编码精确数值的结构化数据规范。我们提出内省与交互式视觉定位(IVG)框架,融合(1)基于规范的内省机制,主动查询底层数据以获取确定性证据;(2)基于视图的交互机制,通过操控可视化界面消除视觉歧义。为避免VLM偏差,我们构建iPlotBench基准,包含500个交互式Plotly图表,共6,706个二元问题及真实数据规范。实验表明,内省显著提升数据重构保真度,结合交互后实现最高问答准确率0.81,在重叠几何结构上相较基线提升6.7%。我们进一步在部署的智能体中验证其能力,实现自主数据探索与实时人机协作。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) frequently misread values, hallucinate details, and confuse overlapping elements in charts. Current approaches rely solely on pixel interpretation, creating a Pixel-Only Bottleneck: agents treat interactive charts as static images, losing access to the structured specification that encodes exact values. We introduce Introspective and Interactive Visual Grounding (IVG), a framework that combines (1) spec-grounded introspection, which queries the underlying specification for deterministic evidence, with (2) view-grounded interaction, which manipulates the view to resolve visual ambiguity. To enable evaluation without VLM bias, we present iPlotBench, a benchmark of 500 interactive Plotly figures with 6,706 binary questions and ground-truth specifications. Experiments show that introspection improves data reconstruction fidelity, while the combination with interaction achieves the highest QA accuracy (0.81), with +6.7 % gains on overlapping geometries. We further demonstrate IVG in deployed agents that explore data autonomously and collaborate with human users in real time.

视觉定位图表理解交互式分析智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。