提出成对分位数回归,用于分析双样本相似度评分的分布特性。
On Pairwise Quantile Regression - Statistical Guarantees and Applications

- 基于配对pinball损失的最小化求解方法
- 理论证明了快速学习速率和泛化误差界
- 适用于人脸识别等双样本相似性分析场景
分位数回归是分析高方差条件下响应变量Y与协变量Z关系的强大工具,超越传统最小二乘回归。本文将该方法拓展至成对设置:目标变量为两个独立观测间的相似度分数(如人脸照片),解释变量为对应观测的协变量(如年龄、发色)。通过研究配对pinball损失的经验最小化解,利用U-过程的精细浓度不等式,建立了该统计学习问题的理论保证,证明了在较弱条件下可实现快速收敛率。模拟实验验证了方法的有效性;应用研究表明,该方法能有效分析人脸识别中相似度评分的误差分布,具有实际应用价值。
原文摘要 · Abstract (English)
Quantile regression provides a powerful tool for summarizing the conditional distribution of a real-valued random variable (r.v.) of interest $Y$ as a function of covariates $Z$ in cases where it shows a large dispersion with high probability, going beyond the situation where standard least square regression is informative/predictive. This article aims to extend this methodology to the pairwise setting, where the variable to be explained is a similarity score between two independent observations (e.g., pixelated ID photos used as input to biometric systems), and the explanatory variables consist of the pair of covariates attached to these observations, such as age or hair color. We establish theoretical guarantees for solutions of this statistical learning problem, considered here as empirical minimizers of a pairwise version of the pinball loss. Leveraging sharp concentration results for $U$-processes, we prove generalization bounds and identify mild conditions under which fast learning rates can be achieved. Confirming the probabilistic analysis, experiments based on simulation data also provide solid empirical evidence of the validity of the methodology promoted here for pairwise quantile regression. Finally, its usefulness from an application perspective is demonstrated by a detailed study aimed at analyzing errors in similarity scoring for facial recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。