梳理高维数据嵌入方法,给出实用指南和评估标准。
Low-dimensional embeddings of high-dimensional data
- 系统分析主流嵌入算法的原理与适用场景。
- 在多个数据集上评估方法性能,提炼出可复用的最佳实践。
- 适合研究人员和工程师快速选择合适嵌入工具。
高维数据在生物、人文等多个领域日益普遍。由于直接处理高维数据存在挑战,对生成低维表示(嵌入)以实现可视化、探索与分析的需求空前增长。近年来涌现大量嵌入算法,应用遍及科研与产业。然而该领域发展迅速且分散,面临技术难题与根本性争议,从业者缺乏有效指导。本文旨在提升研究一致性,提供近期进展的详细批判性综述,推导出创建与使用低维嵌入的最佳实践,对主流方法在多种数据集上的表现进行评估,并讨论该领域现存挑战与开放问题。
原文摘要 · Abstract (English)
Large collections of high-dimensional data have become nearly ubiquitous across many academic fields and application domains, ranging from biology to the humanities. Since working directly with high-dimensional data poses challenges, the demand for algorithms that create low-dimensional representations, or embeddings, for data visualization, exploration, and analysis is now greater than ever. In recent years, numerous embedding algorithms have been developed, and their usage has become widespread in research and industry. This surge of interest has resulted in a large and fragmented research field that faces technical challenges alongside fundamental debates, and it has left practitioners without clear guidance on how to effectively employ existing methods. Aiming to increase coherence and facilitate future work, in this review we provide a detailed and critical overview of recent developments, derive a list of best practices for creating and using low-dimensional embeddings, evaluate popular approaches on a variety of datasets, and discuss the remaining challenges and open problems in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。