梳理可视化生成文本的现状与挑战,构建系统性分类框架。
Natural Language Generation for Visualizations: State of the Art, Challenges and Future Directions
- 提出五类核心问题(Wh-questions)划分生成任务设计空间
- 归纳文本生成输入输出及应用场景,涵盖图表描述与数据叙事
- 为研究者提供可复用的分类体系,助力未来方向探索
自然语言与可视化是人类沟通中互补的两种模态,在有效传递信息方面发挥关键作用。可视化帮助人们发现数据中的趋势、模式和异常,而自然语言描述则有助于解释这些洞察。因此,将文本与可视化结合是一种有效传达数据核心信息的普遍方法。随着自然语言生成(NLG)的发展,自动生成可视化文本描述的研究日益受到关注,可用于图表标题生成、回答图表相关问题或讲述数据驱动的故事。本文系统综述了可视化领域自然语言生成的最新进展,提出一个问题分类体系。该任务属于可视化自然语言接口(NLI)范畴,近年来受到学术界和产业界的广泛关注。为聚焦研究范围,本文主要关注面向可视化文本生成的研究工作。通过提出五个核心问题——为何、如何进行可视化文本生成,输入输出是什么,以及生成文本在何时何地融入可视化——来刻画任务本质与解决方案的设计空间,并据此对已有方法进行分类。最后,讨论该领域面临的关键挑战及未来研究方向。
原文摘要 · Abstract (English)
Natural language and visualization are two complementary modalities of human communication that play a crucial role in conveying information effectively. While visualizations help people discover trends, patterns, and anomalies in data, natural language descriptions help explain these insights. Thus, combining text with visualizations is a prevalent technique for effectively delivering the core message of the data. Given the rise of natural language generation (NLG), there is a growing interest in automatically creating natural language descriptions for visualizations, which can be used as chart captions, answering questions about charts, or telling data-driven stories. In this survey, we systematically review the state of the art on NLG for visualizations and introduce a taxonomy of the problem. The NLG tasks fall within the domain of Natural Language Interfaces (NLI) for visualization, an area that has garnered significant attention from both the research community and industry. To narrow down the scope of the survey, we primarily concentrate on the research works that focus on text generation for visualizations. To characterize the NLG problem and the design space of proposed solutions, we pose five Wh-questions, why and how NLG tasks are performed for visualizations, what the task inputs and outputs are, as well as where and when the generated texts are integrated with visualizations. We categorize the solutions used in the surveyed papers based on these "five Wh-questions." Finally, we discuss the key challenges and potential avenues for future research in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。