通过聚类相关性选择评估指标,减少降维评价中的偏差。
Metric Design != Metric Behavior: Improving Metric Selection for the Unbiased Evaluation of Dimensionality Reduction
- 按指标实际相关性聚类,而非设计初衷
- 选出每类代表性指标,避免重复评价
- 提升降维评估稳定性,适合严谨对比研究
评估降维(DR)投影在保留高维数据结构方面的准确性对可靠的可视化分析至关重要。针对不同结构特征已发展出多种评估指标。然而,若无意中选择高度相关的指标(即测量相似结构特征的指标),会导致评价结果偏倚,偏向于强调这些特征的降维方法。为解决此问题,我们提出一种新工作流程:基于指标的实证相关性而非其设计意图进行聚类,以减少评估偏差。该流程通过计算指标间的成对相关性,聚类指标以最小化重叠,并从每个聚类中选取一个代表性指标。定量实验表明,该方法提升了降维评估的稳定性,证明其有助于缓解评价偏差。
原文摘要 · Abstract (English)
Evaluating the accuracy of dimensionality reduction (DR) projections in preserving the structure of high-dimensional data is crucial for reliable visual analytics. Diverse evaluation metrics targeting different structural characteristics have thus been developed. However, evaluations of DR projections can become biased if highly correlated metrics--those measuring similar structural characteristics--are inadvertently selected, favoring DR techniques that emphasize those characteristics. To address this issue, we propose a novel workflow that reduces bias in the selection of evaluation metrics by clustering metrics based on their empirical correlations rather than on their intended design characteristics alone. Our workflow works by computing metric similarity using pairwise correlations, clustering metrics to minimize overlap, and selecting a representative metric from each cluster. Quantitative experiments demonstrate that our approach improves the stability of DR evaluation, which indicates that our workflow contributes to mitigating evaluation bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。