arXiv:2501.09521cs.HCcs.CL2025-01被引 3

让大模型通过图文结合实现科学数据对话式可视化。

Augmenting a Large Language Model with a Combination of Text and Visual Data for Conversational Visualization of Global Geospatial Data

  • 用图文结合的结构化文本增强大模型,无需微调。
  • 可准确回答基于已渲染可视化图表的问题。
  • 适合需要交互式地理空间数据问答的科研场景。

我们提出一种方法,通过融合文本描述与可视化快照,增强大语言模型(LLM)在科学数据可视化中的问答能力,使对话式可视化成为可能。传统LLM缺乏视觉上下文信息,难以处理可视化交互任务。本方法将可视化图表及其对应文本描述提取为高度紧凑但描述充分的结构化文本文件,有效为LLM提供上下文信息,且无需任何微调。该方法适用于任意已最终渲染的可视化,只要其配有相应的文本说明。

原文摘要 · Abstract (English)

We present a method for augmenting a Large Language Model (LLM) with a combination of text and visual data to enable accurate question answering in visualization of scientific data, making conversational visualization possible. LLMs struggle with tasks like visual data interaction, as they lack contextual visual information. We address this problem by merging a text description of a visualization and dataset with snapshots of the visualization. We extract their essential features into a structured text file, highly compact, yet descriptive enough to appropriately augment the LLM with contextual information, without any fine-tuning. This approach can be applied to any visualization that is already finally rendered, as long as it is associated with some textual description.

对话式可视化大模型地理空间数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。