提出几何化方法融合模糊距离,提升高维数据降维可视化效果
Merging Hazy Sets with m-Schemes: A Geometric Approach to Data Visualization
- 用m-方案统一处理局部密度自适应的相似性函数
- 通过几何构造优化距离嵌入,保留数据拓扑结构
- 适合关注数据可视化与度量学习的研究者
许多机器学习算法试图将高维度量数据在二维空间中可视化,以突出数据的本质几何与拓扑特征。本文提出一种框架,用于聚合源自局部调整度量的相异函数,该调整基于密度感知归一化,如IsUMap方法所采用。我们形式化这些方法为m-方案,这是一类与概率度量中的t-范数和t-反范数密切相关的数学方法,也与信息论中的复合法则相关。m-方案提供了一种灵活且理论严谨的方法,用于改进基于距离的嵌入质量。
原文摘要 · Abstract (English)
Many machine learning algorithms try to visualize high dimensional metric data in 2D in such a way that the essential geometric and topological features of the data are highlighted. In this paper, we introduce a framework for aggregating dissimilarity functions that arise from locally adjusting a metric through density-aware normalization, as employed in the IsUMap method. We formalize these approaches as m-schemes, a class of methods closely related to t-norms and t-conorms in probabilistic metrics, as well as to composition laws in information theory. These m-schemes provide a flexible and theoretically grounded approach to refining distance-based embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。