arXiv:2502.08041cs.LGcs.IT2025-02被引 1

提出分类可判性度量,揭示分类上限不可逾越。

The Art of Misclassification: Too Many Classes, Not Enough Points

  • 用熵量化类别重叠程度,定义分类难易的理论极限
  • 无论模型多强数据多大,准确率无法突破此上限
  • 适合研究分类瓶颈与模型性能边界的研究者

分类是人工智能与机器学习中的核心问题,尽管持续投入大量资源提升模型与数据规模,但分类任务最终受限于数据集的内在特性,而非算力或模型复杂度。本文提出一种基于熵的分类可判性度量,通过评估给定特征表示下类别分配的不确定性,量化分类问题的固有难度。该度量反映类别重叠程度,符合人类直觉,并作为分类性能的理论上限。结果表明,任何分类器在特定问题上都无法突破此上限,无论其架构或数据量如何。本方法为理解分类的本质局限与根本模糊性提供了理论框架。

原文摘要 · Abstract (English)

Classification is a ubiquitous and fundamental problem in artificial intelligence and machine learning, with extensive efforts dedicated to developing more powerful classifiers and larger datasets. However, the classification task is ultimately constrained by the intrinsic properties of datasets, independently of computational power or model complexity. In this work, we introduce a formal entropy-based measure of classificability, which quantifies the inherent difficulty of a classification problem by assessing the uncertainty in class assignments given feature representations. This measure captures the degree of class overlap and aligns with human intuition, serving as an upper bound on classification performance for classification problems. Our results establish a theoretical limit beyond which no classifier can improve the classification accuracy, regardless of the architecture or amount of data, in a given problem. Our approach provides a principled framework for understanding when classification is inherently fallible and fundamentally ambiguous.

分类上限熵度量理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。