让大模型像专家一样一步步分析数据,生成更高质量的可视化图表。
InsightChain: Optimized Chain-of-Insight Analytics for LLM-driven Data Visualization

- 分四步引导模型逐步思考:探索、聚焦、测试、呈现。
- 在十个领域中表现优于现有方法,跨领域效果也稳定提升。
- 新评估指标融合文本与视觉判断,更贴近真实分析过程。
大型语言模型(LLMs)正被广泛用于自动化数据可视化,但现有方法常将可视化生成视为从用户查询到图表或代码的单步映射,忽略了专家分析师的迭代式分析过程。本文提出InsightChain,一个四阶段可视化提示流程(探索—聚焦—测试—呈现),模拟专家分析工作流,并引入VG-COPRO,一种适应多阶段可执行流程的视觉引导自动提示优化(APO)方法。为弥补复杂可视化评估的不足,我们设计了洞察进展度量(IPM),结合四个文本维度与一个视觉维度构成评分体系。通过100条链的人类预实验和覆盖全部十领域的300条链代理评估验证,结果表明InsightChain在多个公共数据集上持续优于竞争基线。现有APO方法在该多阶段任务中未能带来一致提升,而VG-COPRO在同域与跨域场景中均显著改善性能。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for automated data visualization, yet existing approaches often frame visualization generation as a single-step mapping from user query to figure or code, overlooking the iterative analytical reasoning process of expert analysts. We present InsightChain, a four-stage visualization prompting pipeline (Explore--Focus--Test--Present) that emulates expert analytical workflows, together with VG-COPRO, a vision-guided automatic prompt optimization (APO) method adapted to jointly optimize such multi-stage, executable pipelines. To address the evaluation gap for complex data visualization, we introduce the Insight Progression Metric (IPM), a rubric combining four text-based dimensions with a vision-based dimension. We assess IPM through a 100-chain human pilot and an expanded 300-chain agent-based evaluation spanning all ten domains. Experiments on public datasets show that InsightChain consistently outperforms competing prompting baselines. Existing APO methods fail to yield consistent gains on this multi-stage task, whereas VG-COPRO improves performance in both in-domain and cross-domain settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。