NOMAD投影可高效可视化超大规模数据,突破传统方法瓶颈。
NOMAD Projection
- 基于负相关与均值亲和力的非线性降维新方法
- 在多卡训练下实现毫秒级处理,性能超越现有最优方法
- 适合处理百万级数据集,尤其适用于跨语言数据探索
生成式AI的快速普及导致人工智能模型所消耗和生成的数据集规模急剧增长。传统的无结构数据可视化方法(如t-SNE和UMAP)未能跟上数据量扩展的步伐,给AI可解释性带来了挑战,因为其依赖这些方法进行探索性数据分析。本文提出负相关或均值亲和力判别(NOMAD)投影,这是首个可在训练时使用多块GPU运行的非线性降维无结构数据可视化方法。我们提供了理论依据,表明NOMAD投影是InfoNC-t-SNE损失的一个近似上界,并通过实验证明其在性能和速度上均优于现有最先进方法。我们展示了该方法的可扩展性,首次完成了多语言维基百科的完整数据映射。
原文摘要 · Abstract (English)
The rapid adoption of generative AI has driven an explosion in the size of datasets consumed and produced by AI models. Traditional methods for unstructured data visualization, such as t-SNE and UMAP, have not kept up with the pace of dataset scaling. This presents a significant challenge for AI explainability, which relies on methods such as t-SNE and UMAP for exploratory data analysis. In this paper, we introduce Negative Or Mean Affinity Discrimination (NOMAD) Projection, the first method for unstructured data visualization via nonlinear dimensionality reduction that can run on multiple GPUs at train time. We provide theory that situates NOMAD Projection as an approximate upper bound on the InfoNC-t-SNE loss, and empirical results that demonstrate NOMAD Projection's superior performance and speed profile compared to existing state-of-the-art methods. We demonstrate the scalability of NOMAD Projection by computing the first complete data map of Multilingual Wikipedia.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。