arXiv:2411.18657cs.AIcs.HC2024-11KDD被引 1

用强化学习精简可视化推荐的统计量,提速10倍且误差小。

ScaleViz: Scaling Visualization Recommendation Models on Large Data

  • 基于强化学习动态选择最优统计量组合,节省计算时间。
  • 在三个真实大数据集上,速度提升10倍,误差极低。
  • 适合需要快速分析大规模数据的科研与工程人员。

自动化可视化推荐(vis-rec)帮助用户从新数据集中提取关键洞察。传统方法先计算大量统计数据,再用机器学习模型评分或分类多种可视化选项,以推荐最有效的方案。然而,现有先进模型依赖海量昂贵的统计数据,导致在大规模数据上运行时计算时间过长,难以应用于大多数现实世界中的复杂大型数据集。本文提出一种基于强化学习(RL)的新框架,给定一个可视化推荐模型和用户设定的时间预算,自动识别出在限定时间内生成有效视觉洞察的最佳输入统计量集合。我们在两个最先进的可视化推荐模型上,针对三个真实大型数据集进行了实验,结果表明该方法能显著缩短可视化时间,引入的误差极小。相比基线方法,在相似误差水平下,本方法提速约10倍。

原文摘要 · Abstract (English)

Automated visualization recommendations (vis-rec) help users to derive crucial insights from new datasets. Typically, such automated vis-rec models first calculate a large number of statistics from the datasets and then use machine-learning models to score or classify multiple visualizations choices to recommend the most effective ones, as per the statistics. However, state-of-the art models rely on very large number of expensive statistics and therefore using such models on large datasets become infeasible due to prohibitively large computational time, limiting the effectiveness of such techniques to most real world complex and large datasets. In this paper, we propose a novel reinforcement-learning (RL) based framework that takes a given vis-rec model and a time-budget from the user and identifies the best set of input statistics that would be most effective while generating the visual insights within a given time budget, using the given model. Using two state-of-the-art vis-rec models applied on three large real-world datasets, we show the effectiveness of our technique in significantly reducing time-to visualize with very small amount of introduced error. Our approach is about 10X times faster compared to the baseline approaches that introduce similar amounts of error.

可视化推荐强化学习高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。