arXiv:2411.07983cs.AI2024-11被引 6

用吉尼系数统一评估向量空间中多对多相似性,提升模型训练效果。

Gini Coefficient as a Unified Metric for Evaluating Many-versus-Many Similarity in Vector Spaces

  • 以吉尼系数衡量向量间全局相似性,无需预设类别。
  • 高吉尼系数样本组内相似度更高,模型性能显著优于随机采样。
  • 适合数据稀疏场景下筛选代表性训练样本,尤其适用于小样本学习。

我们证明吉尼系数可作为统一指标,用于评估向量空间中的多对多(全对全)相似性。对多种图像数据集的分析表明,吉尼系数最高的图像彼此间最相似,而吉尼系数最低的图像则最不相似。该关系在不同语料库的文本嵌入向量中同样成立,验证了方法的一致性与跨数据类型的普适性。此外,我们发现选择与测试集分布匹配的训练样本,比保证数据多样性更为关键。具有较高吉尼系数的典型、代表性训练样本,其模型性能远超仅具多样性的低吉尼系数数据集。因此,吉尼系数可有效指导机器学习样本选择,在信息极度稀疏的条件下,本方法优于随机采样。

原文摘要 · Abstract (English)

We demonstrate that Gini coefficients can be used as unified metrics to evaluate many-versus-many (all-to-all) similarity in vector spaces. Our analysis of various image datasets shows that images with the highest Gini coefficients tend to be the most similar to one another, while images with the lowest Gini coefficients are the least similar. We also show that this relationship holds true for vectorized text embeddings from various corpuses, highlighting the consistency of our method and its broad applicability across different types of data. Additionally, we demonstrate that selecting machine learning training samples that closely match the distribution of the testing dataset is far more important than ensuring data diversity. Selection of exemplary and iconic training samples with higher Gini coefficients leads to significantly better model performance compared to simply having a diverse training set with lower Gini coefficients. Thus, Gini coefficients can serve as effective criteria for selecting machine learning training samples, with our selection method outperforming random sampling methods in very sparse information settings.

相似性评估吉尼系数样本选择小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。