提出可逆降维方法FloDR,保留高维信息并支持精确反推与可视化诊断。
FloDR: An invertible dimensionality reduction method based on a normalising flow

- 基于可逆归一化流构建降维模型,保留全部高维坐标。
- 提供精确逆映射与密度估计,支持真实数据重构与置信度检验。
- 可生成条件扩散与隐藏对比场,揭示嵌入中被丢弃的信息量。
高维数据的二维嵌入常被误读,因t-SNE、UMAP等方法在优化过程中丢弃了支撑距离、簇间关系及结构信息的关键内容。本文提出FloDR,一种基于可逆归一化流的降维方法。尽管仅使用前两个输出坐标生成二维嵌入,但完整保留其余坐标。训练后模型具备精确逆映射与密度估计,可用于基于真实逆映射的诊断性可视化。具体而言,我们绘制两个场:条件扩散(衡量每个嵌入位置在输入空间中的不确定性),以及隐藏对比(衡量所选两坐标对标签对比信息的丢失程度)。两个场均通过预留数据集和自助法置信度进行预设检验,若不通过则标记为拒绝。
原文摘要 · Abstract (English)
It is common for two-dimensional embeddings of high-dimensional data to be read far beyond what they can support. Distances in and between clusters, the meaning behind empty spaces, and the amount of structure hidden at each point are generally invisible in the output of methods such as t-SNE and UMAP. This is because the information that could support the meaning of these properties is discarded during the optimisation process. Here, we present FloDR, a dimensionality reduction method that embeds data through an invertible normalising flow. While FloDR only uses the first two output coordinates to create a two-dimensional embedding, it retains the remaining coordinates rather than discarding them. In addition to the embedding, an exact inverse and an exact density are properties of a trained mapping, which enable diagnostic visualisations that are computed from the exact inverse of the model that drew the layout rather than from an approximate one. Specifically, we draw two fields, the conditional spread, which measures how much of the original data remains undetermined at each embedding position in input units, and the hidden contrast, which measures how much information about a labelled contrast the two plotted coordinates discard. Both fields are rendered with a prespecified test against a held out portion of the input data and a bootstrap confidence. A field that fails the test is reported as refused.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。