用拓扑方法分析数据几何结构,预测训练质量。
Predict Training Data Quality via Its Geometry in Metric Space
- 通过持久同调提取度量空间中的数据拓扑特征
- 发现数据几何结构与模型性能强相关
- 适合关注数据质量评估的研究者
高质量训练数据是机器学习与人工智能的基础,决定模型的学习方式与表现。尽管已知某些类型的数据对训练更有效,但数据几何结构对模型性能的影响仍鲜被研究。我们提出,数据表示的丰富性与内部冗余的消除对学习结果至关重要。为此,我们采用持久同调(persistent homology)从度量空间中提取数据的拓扑特征,提供了一种超越熵基度量的多样性量化方法。研究结果表明,持久同调是分析和提升驱动人工智能系统的训练数据的有效工具。
原文摘要 · Abstract (English)
High-quality training data is the foundation of machine learning and artificial intelligence, shaping how models learn and perform. Although much is known about what types of data are effective for training, the impact of the data's geometric structure on model performance remains largely underexplored. We propose that both the richness of representation and the elimination of redundancy within training data critically influence learning outcomes. To investigate this, we employ persistent homology to extract topological features from data within a metric space, thereby offering a principled way to quantify diversity beyond entropy-based measures. Our findings highlight persistent homology as a powerful tool for analyzing and enhancing the training data that drives AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。