验证手写数字MNIST数据集是否线性可分
On Linear Separability of the MNIST Handwritten Digits Dataset
- 对比分析训练、测试及合并集的成对与多类线性分离性
- 发现测试集在多数情况下无法实现完全线性可分
- 为机器学习模型选择提供理论依据,适合研究分类边界者
MNIST手写数字数据集虽规模小、分辨率低,仍是模式识别与图像分类模型的重要基准。线性可分性是众多统计与机器学习方法的核心概念。尽管该数据集历史悠久,其是否线性可分却始终缺乏定论,学术与非学术来源存在相互矛盾的说法。本文通过系统实证研究,分别考察训练集、测试集及合并集在成对分类与一对其余分类下的线性可分性,回顾理论评估方法与前沿工具,全面分析所有相关组合,报告具体结果。
原文摘要 · Abstract (English)
The MNIST dataset containing thousands of handwritten digit images is still a fundamental benchmark for evaluating various pattern-recognition and image-classification models. Linear separability is a key concept in many statistical and machine-learning techniques. Despite the long history of the MNIST dataset and its relative simplicity in size and resolution, the question of whether the dataset is linearly separable has never been fully answered -- scientific and informal sources share conflicting claims. This paper aims to provide a comprehensive empirical investigation to address this question, distinguishing pairwise and one-vs-rest separation of the training, the test and the combined sets, respectively. It reviews the theoretical approaches to assessing linear separability, alongside state-of-the-art methods and tools, then systematically examines all relevant assemblies, and reports the findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。