arXiv:2605.28767cs.LGstat.ML2026-05被引 7

针对多标签分类的复杂评估指标,提出可证明收敛的优化算法。

Principled Algorithms for Optimizing Generalized Metrics in Multi-Label Learning

  • 设计新型代理损失函数,实现非渐近保证的理论一致性
  • 算法时间复杂度严格为O(l),无需近似计算
  • 适用于高稀疏、大规模数据,对齐真实评估指标

许多现实世界的分类任务需要对每个样本预测多个标签,这要求优化复杂的评估指标,如F- measure和杰卡德指数。尽管经验效用最大化(EUM)框架适用于这些总体层面的指标,但现有理论结果大多局限于渐近贝叶斯一致性。本文在EUM框架下,针对广义度量开发了有原则的学习算法,基于更强的H-一致性概念。关键贡献是设计了新型的多标签学习代理损失函数,具备可证明的H-一致性界,支持针对假设类和有限样本的非渐近优化保证。重要的是,我们证明这些组合形式的代理函数可精确分解,计算时间严格为O(l),无需近似。在此基础上,提出MMO(多标签度量优化)算法族,用于优化广义线性分数型度量。通过大量实验验证,该方法在大规模数据集(MS-COCO、Reuters-21578)上展现出稳健的可扩展性和优于当前最优连续基线的性能,在高稀疏、深度学习场景中表现突出。结果兼具理论严谨性与实际有效性。

原文摘要 · Abstract (English)

Many real-world classification tasks require predicting multiple labels per instance, necessitating the optimization of complex evaluation metrics such as the $F$-measure and Jaccard index. While the Empirical Utility Maximization (EUM) framework is natural for these population-level metrics, existing theoretical results are largely limited to asymptotic Bayes-consistency. In this paper, we develop principled learning algorithms for optimizing a broad class of generalized metrics within the EUM framework, grounded in the stronger notion of $H$-consistency. Our key contribution is the design of novel surrogate loss functions for multi-label learning that admit provable $H$-consistency bounds, enabling optimization with non-asymptotic guarantees tailored to the hypothesis class and finite samples. Crucially, we prove these combinatorially formulated surrogates decompose exactly, operating in strictly $O(l)$ time without approximations. Building on this foundation, we introduce MMO (Multi-Label Metric Optimization), a new family of algorithms for optimizing generalized linear-fractional metrics. We validate our approach through extensive experiments, demonstrating robust scalability and superior performance over state-of-the-art continuous baselines on large-scale datasets (MS-COCO, Reuters-21578) in high-sparsity, deep learning regimes. Our results offer both theoretical rigor and practical effectiveness for general multi-label metric optimization.

多标签学习优化算法理论保证评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。