arXiv:2512.08371cs.LGstat.ML2025-12被引 2

针对多标签数据中稀有标签样本不足问题,提出一种考虑标签依赖关系的加权采样方法。

A Multivariate Bernoulli-Based Sampling Method for Multi-Label Data with Application to Meta-Research

  • 基于多元伯努利分布建模标签依赖,动态计算每种标签组合的权重
  • 在64个生物医学主题类别数据上,显著降低最常见与最罕见类别的频率差异
  • 适合需要均衡表示稀有标签的元研究、文献分析等场景

多标签数据中若标签不互斥且频次差异大,难以获取既能充分覆盖稀有标签、又保持已知分布特征的样本。本文将多标签问题建模为多元伯努利分布,提出一种新采样算法,利用观测标签频率估计分布参数,并为每种标签组合计算权重,使加权采样既符合目标分布特征,又考虑标签间依赖关系。该方法应用于包含64个生物医学主题类别的Web of Science文章样本,旨在保持类别频率排序,缩小最常见与最罕见类别间的频率差距,并捕捉类别依赖。实验结果表明,该方法生成了更均衡的子样本,提升了少数类别代表性。

原文摘要 · Abstract (English)

Datasets may contain observations with multiple labels. If the labels are not mutually exclusive, and if the labels vary greatly in frequency, obtaining a sample that includes sufficient observations with scarcer labels to make inferences about those labels, and which deviates from the population frequencies in a known manner, creates challenges. In this paper, we consider a multivariate Bernoulli distribution as our underlying distribution of a multi-label problem. We present a novel sampling algorithm that takes label dependencies into account. It uses observed label frequencies to estimate multivariate Bernoulli distribution parameters and calculates weights for each label combination. This approach ensures the weighted sampling acquires target distribution characteristics while accounting for label dependencies. We applied this approach to a variety of datasets, including a sample of research articles from Web of Science labeled with 64 biomedical topic categories. We aimed to preserve category frequency order, reduce frequency differences between most and least common categories, and account for category dependencies. This approach produced a more balanced sub-sample, enhancing the representation of minority categories.

多标签采样伯努利分布数据均衡元研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。