无需训练即可快速评估人脸数据集的潜在性能
Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets

- 通过邻居一致性与嵌入空间复杂度综合评估数据质量
- 仅用小样本或轻量模型即可预测大模型表现
- 适合数据清洗、筛选和资源优化场景
我们提出一种名为内在质量(Intrinsic Quality, IQ)的无验证评估指标,可无需完整训练即估计人脸识别(FR)数据集产生高性能模型的潜力。IQ结合两项指标:(i) 邻居一致性得分,通过最近邻的标签一致性量化局部身份一致性;(ii) 全局表示子空间复杂度(有效秩,ER),刻画嵌入空间的几何结构与数据多样性。该方法仅需轻量代理模型或数据子集即可实现快速评估,适用于大规模数据集的诊断与清洗。我们设计了针对干净、噪声及混合质量数据集的实验协议,并建立评估方法以验证IQ对下游性能的预测能力。
原文摘要 · Abstract (English)
We propose Intrinsic Quality (IQ), a validation-free metric designed to estimate the inherent potential of face recognition (FR) datasets to produce high-performance models without the need for full-scale training. IQ integrates two components: (i) a Neighbor-Consistency Score that quantifies local identity label agreement via nearest neighbors, and (ii) Global Representation Subspace Complexity (Effective Rank, ER), which captures the underlying embedding geometry and dataset diversity. IQ allows for rapid evaluation using lightweight proxy models or data subsets, facilitating dataset diagnosis and curation prior to resource-intensive full-scale training. We describe an experimental protocol tailored to clean, noisy, and mixed-quality FR datasets, and outline evaluation methodologies to validate IQ's predictive power for downstream performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。