arXiv:2508.18951cs.LG2025-08

比较三种模型估算多标签数据条件协方差,发现多重正态模型误差最小。

Estimating Conditional Covariance between labels for Multilabel Data

  • 用多重正态、伯努利和分阶段逻辑回归模型分析标签条件协方差
  • 所有模型在强协方差下表现相当,但均误判恒定协方差为依赖协方差
  • 多重正态模型误差最低,适合需要精确协方差估计的任务

多标签数据在应用模型前应分析标签间依赖关系。由于标签值受协变量向量 $\\(vec{x}$ 影响,无法直接测量标签独立性,需通过多元正态模型检验条件标签协方差。然而,该模型仅提供其共现分布的协方差估计,可能在估计恒定与依赖协方差时不可靠。本文对比了三种模型(多元正态、多元伯努利、分阶段逻辑回归)在估计恒定与依赖协方差上的表现。实验显示,所有模型在协方差强度较大时性能相近,但均错误检测出恒定协方差数据中存在依赖协方差。其中,多元正态模型误差率最低。

原文摘要 · Abstract (English)

Multilabel data should be analysed for label dependence before applying multilabel models. Independence between multilabel data labels cannot be measured directly from the label values due to their dependence on the set of covariates $\vec{x}$, but can be measured by examining the conditional label covariance using a multivariate Probit model. Unfortunately, the multivariate Probit model provides an estimate of its copula covariance, and so might not be reliable in estimating constant covariance and dependent covariance. In this article, we compare three models (Multivariate Probit, Multivariate Bernoulli and Staged Logit) for estimating the constant and dependent multilabel conditional label covariance. We provide an experiment that allows us to observe each model's measurement of conditional covariance. We found that all models measure constant and dependent covariance equally well, depending on the strength of the covariance, but the models all falsely detect that dependent covariance is present for data where constant covariance is present. Of the three models, the Multivariate Probit model had the lowest error rate.

多标签学习协方差估计概率模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。