提出一种衡量数据集相似性的几何度量方法,能有效捕捉高维数据结构差异。
Magnitude Distance: A Geometric Measure of Dataset Similarity
- 基于度量空间的幅度概念定义新距离,可调参数控制对全局与细节的敏感度
- 在高维场景下仍具区分能力,理论证明其在不同尺度下的收敛行为
- 可作为生成模型训练目标,实验显示效果媲美传统基于距离的方法
量化数据集间的距离是数学和机器学习中的基本问题。我们提出一种新的距离度量——幅度距离,该度量基于度量空间的幅度概念,适用于有限数据集。该距离包含一个可调缩放参数 $t$,用于控制对全局结构(小 $t$)和精细细节(大 $t$)的敏感性。我们证明了幅度距离的若干理论性质,包括其在不同尺度下的极限行为,以及满足关键度量性质的条件。与经典距离相比,当尺度适当调节时,幅度距离在高维设置中仍保持区分能力。此外,我们展示了如何将幅度距离用作推前生成模型的训练目标。实验结果支持我们的理论分析,并表明幅度距离提供了有意义的信号,其表现可与现有基于距离的生成方法相媲美。
原文摘要 · Abstract (English)
Quantifying the distance between datasets is a fundamental question in mathematics and machine learning. We propose \textit{magnitude distance}, a novel distance metric defined on finite datasets using the notion of the \emph{magnitude} of a metric space. The proposed distance incorporates a tunable scaling parameter, $t$, that controls the sensitivity to global structure (small $t$) and finer details (large $t$). We prove several theoretical properties of magnitude distance, including its limiting behavior across scales and conditions under which it satisfies key metric properties. In contrast to classical distances, we show that magnitude distance remains discriminative in high-dimensional settings when the scale is appropriately tuned. We further demonstrate how magnitude distance can be used as a training objective for push-forward generative models. Our experimental results support our theoretical analysis and demonstrate that magnitude distance provides meaningful signals, comparable to established distance-based generative approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。