为多维尺度分析提供置信集,可量化嵌入结果的不确定性。
Confidence Sets for Multidimensional Scaling
- 基于统计理论构建了噪声数据下嵌入结果的置信集。
- 乘子自举法能自动适应异方差噪声,提升小样本精度。
- 适用于需要评估嵌入可靠性的人工智能与数据分析场景。
我们为应用于噪声相异性数据的经典多维尺度分析(CMDS)建立了正式的统计框架。针对多种噪声模型,我们推导出嵌入结果的分布收敛性,从而可在刚性变换不变意义下构造真正的统一置信集。进一步提出基于自举的置信集构造方法,并提供了有效性理论保证。研究发现,乘子自举法能自动适应异方差噪声(如乘性噪声),而经验自举法则需同方差假设。在有效前提下,两类自举均显著提升有限样本下的准确性。通过数值实验验证了所提方法的实证性能。
原文摘要 · Abstract (English)
We develop a formal statistical framework for classical multidimensional scaling (CMDS) applied to noisy dissimilarity data. We establish distributional convergence results for the embeddings produced by CMDS for various noise models, which enable the construction of \emph{bona~fide} uniform confidence sets for the latent configuration, up to rigid transformations. We further propose bootstrap procedures for constructing these confidence sets and provide theoretical guarantees for their validity. We find that the multiplier bootstrap adapts automatically to heteroscedastic noise such as multiplicative noise, while the empirical bootstrap seems to require homoscedasticity. Either form of bootstrap, when valid, is shown to substantially improve finite-sample accuracy. The empirical performance of the proposed methods is demonstrated through numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。