arXiv:2506.16571cs.HCcs.CL2025-06被引 4

用学生设计笔记构建可视化决策理由数据集。

Capturing Visualization Design Rationale

  • 基于学生课程笔记中的真实可视化与设计说明,构建自然语言数据集。
  • 利用大模型生成并验证问答-理由三元组,提炼设计逻辑。
  • 适合研究可视化教学、人机交互与可解释性设计的学者。

现有数据集多聚焦于可视化理解任务,依赖人工构造的可视化与问题,侧重解码而非编码。本文提出新数据集与方法,通过学生在数据可视化课程中撰写的具有设计说明的可视化笔记,捕捉真实的可视化设计理由。这些笔记结合视觉图表与设计阐述,明确展示设计选择背后的逻辑。我们利用大语言模型从笔记内容中自动生成并分类问答-理由三元组,并经过严格验证与筛选,最终构建一个包含可视化设计决策及其对应理由的高质量数据集。

原文摘要 · Abstract (English)

Prior natural language datasets for data visualization have focused on tasks such as visualization literacy assessment, insight generation, and visualization generation from natural language instructions. These studies often rely on controlled setups with purpose-built visualizations and artificially constructed questions. As a result, they tend to prioritize the interpretation of visualizations, focusing on decoding visualizations rather than understanding their encoding. In this paper, we present a new dataset and methodology for probing visualization design rationale through natural language. We leverage a unique source of real-world visualizations and natural language narratives: literate visualization notebooks created by students as part of a data visualization course. These notebooks combine visual artifacts with design exposition, in which students make explicit the rationale behind their design decisions. We also use large language models (LLMs) to generate and categorize question-answer-rationale triples from the narratives and articulations in the notebooks. We then carefully validate the triples and curate a dataset that captures and distills the visualization design choices and corresponding rationales of the students.

可视化设计理由大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。