arXiv:2607.08579cs.HCcs.LG2026-07

可视化工具助你诊断缺失数据并对比不同填补方法效果

ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods

论文配图:ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods
图 1 · 摘自论文原文
  • 整合MICE、随机森林等多类填补方法,支持交互式配置与评估
  • 引入地理加权kNN,融合空间与社会经济距离,追踪数据来源
  • 通过分布重叠、误差指标等视图,直观比较方法差异与敏感变量

缺失数据是科学、社会科学和公共卫生研究中的长期难题,常导致分析偏差,并使分析师需对处理方式负责。本文提出ImputeViz,一个集成的可视化分析仪表盘,支持缺失性诊断、填补模型配置与结果评估。系统整合了MICE、随机森林、XGBoost和kNN等常用方法,通过热力图、共缺失统计与分布诊断等协同视图,揭示缺失模式(MCAR/MAR)及非随机缺失(MNAR)线索。针对地理空间分析需求,提出gKNN——一种融合社会经济与空间距离的改进kNN算法,可显式展示各区域对估算值的贡献,实现基于溯源的可视化问责。用户可通过分布叠加、方法对比摘要(含MAE、RMSE、Delta RMSE与运行时间)及变量级差异视图,跨方法比较性能。缓存结果与锁定坐标轴范围降低切换方法时的认知负担。案例研究显示,该工具能帮助选择有效策略、识别敏感变量并评估模型鲁棒性。

原文摘要 · Abstract (English)

Missing data is a persistent obstacle in scientific, social science, and public health research, often biasing analyses and placing accountability on analysts for how they handle missing values. We introduce ImputeViz, an integrated visual analytics dashboard that supports diagnosing missingness, configuring imputation models, and evaluating results. The system brings together widely used methods, including MICE, Random Forest, XGBoost, and kNN, within an interactive environment that makes missingness patterns explicit. To support geospatial reasoning, we introduce gKNN, a geographically informed kNN variant that blends socioeconomic and spatial distances and exposes donor contributions, enabling provenance-based visual accountability by showing which regions drive each estimate. Our primary contribution is a method-agnostic visual analytics environment that makes cross-method comparison a first-class visual task and integrates gKNN alongside standard methods. Coordinated views reveal missingness structure through heatmaps, co-missingness summaries, and distributional diagnostics that help analysts reason about missingness patterns (MCAR/MAR) and cases where missingness may be non-random (MNAR). Users can compare and tune models and interrogate results via distributional overlays, a Method Comparison Summary reporting MAE, RMSE, Delta RMSE, and runtime for each algorithm on the current target and mask, along with variable-level discrepancy views. Cached per-method results and locked axis scales reduce cognitive overhead from shifting ranges during method switching. These comparisons highlight where methods disagree, which variables are sensitive, and how imputation choices affect downstream summaries. Case studies demonstrate how ImputeViz helps analysts select effective strategies, surface sensitive variables, and assess model robustness.

数据填补可视化分析缺失数据地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。