arXiv:2601.15560cs.CV2026-01

提出新指标RCA,精准检测K-pop人脸生成中的身份错乱问题

Relative Classification Accuracy: A Calibrated Metric for Identity Consistency in Fine-Grained K-pop Face Generation

  • 用基准分类器归一化,提出相对分类准确率RCA评估身份一致性
  • 模型FID仅8.93但RCA低至0.27,暴露严重语义模式崩溃
  • 适合关注生成模型身份准确性的研究人员参考

去噪扩散概率模型(DDPM)在高保真图像生成中取得显著进展,但在细粒度单领域任务中评估其语义可控性仍具挑战。标准指标如FID和Inception Score(IS)难以检测此类场景下的身份错位问题。本文研究用于K-pop偶像人脸生成(32x32)的类条件DDPM,该领域具有高度类间相似性。我们提出校准指标相对分类准确率(RCA),将生成性能相对于一个基准分类器进行归一化。评估发现关键权衡:尽管模型视觉质量高(FID 8.93),却存在严重语义模式崩溃(RCA 0.27),尤其在视觉模糊的身份上表现更差。通过混淆矩阵分析失败模式,归因于分辨率限制与同性别内在相似性。本框架为验证条件生成模型的身份一致性提供了严格标准。

原文摘要 · Abstract (English)

Denoising Diffusion Probabilistic Models (DDPMs) have achieved remarkable success in high-fidelity image generation. However, evaluating their semantic controllability-specifically for fine-grained, single-domain tasks-remains challenging. Standard metrics like FID and Inception Score (IS) often fail to detect identity misalignment in such specialized contexts. In this work, we investigate Class-Conditional DDPMs for K-pop idol face generation (32x32), a domain characterized by high inter-class similarity. We propose a calibrated metric, Relative Classification Accuracy (RCA), which normalizes generative performance against an oracle classifier's baseline. Our evaluation reveals a critical trade-off: while the model achieves high visual quality (FID 8.93), it suffers from severe semantic mode collapse (RCA 0.27), particularly for visually ambiguous identities. We analyze these failure modes through confusion matrices and attribute them to resolution constraints and intra-gender ambiguity. Our framework provides a rigorous standard for verifying identity consistency in conditional generative models.

扩散模型身份一致生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。