arXiv:2602.21572stat.MLcs.LG2026-02

提出新检验方法,精准判断潜类别模型应分几类。

Goodness-of-Fit Tests for Latent Class Models with Ordinal Categorical Data

  • 用归一化残差矩阵最大奇异值做检验统计量
  • 样本量足够时,正确类别数下统计量趋近零
  • 适合心理学、教育学等使用有序分类数据的研究者

有序分类数据在心理学、教育学等社会科学中广泛存在,常见于问卷、测评与调查。潜类别模型可通过响应模式将个体划分为同质类群,揭示未观测异质性。其核心挑战在于确定潜类别数量,该值未知且需从数据中推断。本文提出一种检验统计量,以样本量调整后的归一化残差矩阵最大奇异值为中心。在候选类别数正确的零假设下,其上界以概率收敛至零;在拟合不足的备择假设下,统计量超过固定正数的概率趋于1。该统计量呈现显著二元行为,由此衍生出两种顺序检验算法,可一致估计真实类别数。大量实验验证了理论结果,表明该方法在确定潜类别数量上具有高准确性和可靠性。

原文摘要 · Abstract (English)

Ordinal categorical data are widely collected in psychology, education, and other social sciences, appearing commonly in questionnaires, assessments, and surveys. Latent class models provide a flexible framework for uncovering unobserved heterogeneity by grouping individuals into homogeneous classes based on their response patterns. A fundamental challenge in applying these models is determining the number of latent classes, which is unknown and must be inferred from data. In this paper, we propose one test statistic for this problem. The test statistic centers the largest singular value of a normalized residual matrix by a simple sample-size adjustment. Under the null hypothesis that the candidate number of latent classes is correct, its upper bound converges to zero in probability. Under an under-fitted alternative, the statistic itself exceeds a fixed positive constant with probability approaching one. This sharp dichotomous behavior of the test statistic yields two sequential testing algorithms that consistently estimate the true number of latent classes. Extensive experimental studies confirm the theoretical findings and demonstrate their accuracy and reliability in determining the number of latent classes.

潜类别模型有序数据统计检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。