通过分析200+篇论文,揭示人类如何用可视化向机器学习流程注入知识。
Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics

- 从机器学习、可视化、交互和行为四方面构建编码体系
- 发现人类知识通过交互式可视化以多种路径融入模型构建
- 提出模型构建与信息成本收益分析双理论框架,适合研究人机协同者
可视化分析(VA)在支持机器学习(ML)工作流中作用日益重要,相关方法被称为VIS4ML。尽管模型多自动学习,但工作流仍需大量人工输入,如数据标注、特征工程、架构设计与超参数调优等。本研究调研了过去十年IEEE VIS会议中的200余篇VIS4ML论文,构建了涵盖机器学习特性、可视化、交互与行为四个维度的编码方案。通过对编码数据的分析,揭示了人类知识通过交互式可视化传递至ML工作流的不同路径。基于此,提出将VA视为模型构建过程,并用信息论的成本-收益分析解释其优化工作流的作用。研究为使用可视化分析提升机器学习提供了明确证据。完整论文列表及分析结果见https://vis4ml4hd.github.io/ml-knowledge-inject-va/。
原文摘要 · Abstract (English)
Visual analytics (VA) plays an increasingly important role in supporting machine learning (ML) workflows. In the field of visualization, such approaches and techniques are referred to as VIS4ML. While ML models are mostly learned automatically, the corresponding ML workflows receive a variety of human inputs, such as data labelling, feature engineering, model architecture designing, hyper-parameter tuning, and so on. In this work, we surveyed over 200 VIS4ML papers to gain an understanding of how humans inject their knowledge into ML workflows through interactive visualization. We collected a corpus of VIS4ML papers from the IEEE VIS conferences in the past decade. We developed a coding scheme to facilitate the literature research from four perspectives: characteristics of ML, visualization, interaction, and actions. The analysis of the coded dataset allows us to observe different pathways that transfer human knowledge to ML workflows via interactive visualization. Building on the analysis, we explain the phenomena of VIS4ML using the conceptual model that views VA as model building and the information-theoretic cost-benefit analysis that reasons VA as for optimizing ML workflows. This work provides unequivocal evidence showing the merits of using VA in ML workflows. The full list of surveyed papers, along with all analysis results and figures, is available at https://vis4ml4hd.github.io/ml-knowledge-inject-va/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。