仅用一张样本即可精准分类高度相似文档,像人一样学习概念。
Coordinate Matrix Machine: A Human-level Concept Learning to Classify Very Similar Documents
- 通过捕捉文档结构特征实现人类级概念学习
- 单样本分类准确率超越传统模型和深度学习方法
- 适合资源受限场景,解释性强且环保节能
人类通常只需一个例子就能掌握新概念,而机器学习模型往往需要数百个样本。本文提出坐标矩阵机(CM²),一种专为小规模设计的绿色AI模型,通过识别文档中人类会关注的结构化关键特征,实现对高度相似文档的分类。该模型不依赖大规模预训练和高性能硬件,仅需每类一个样本即可完成学习。实验表明,其性能优于传统向量化方法及复杂的深度学习模型,具备高精度、几何与结构感知能力、可解释性(透明模型)、低延迟、抗类别不平衡、计算成本低等优势,适用于仅靠CPU运行的环境,兼具可持续性与经济可行性。
原文摘要 · Abstract (English)
Human-level concept learning argues that humans typically learn new concepts from a single example, whereas machine learning algorithms typically require hundreds of samples to learn a single concept. Our brain subconsciously identifies important features and learns more effectively. Contribution: In this paper, we present the Coordinate Matrix Machine (CM$^2$). This purpose-built small model augments human intelligence by learning document structures and using this information to classify documents. While modern "Red AI" trends rely on massive pre-training and energy-intensive GPU infrastructure, CM$^2$ is designed as a Green AI solution. It achieves human-level concept learning by identifying only the structural "important features" a human would consider, allowing it to classify very similar documents using only one sample per class. Advantage: Our algorithm outperforms traditional vectorizers and complex deep learning models that require larger datasets and significant compute. By focusing on structural coordinates rather than exhaustive semantic vectors, CM$^2$ offers: 1. High accuracy with minimal data (one-shot learning) 2. Geometric and structural intelligence 3. Green AI and environmental sustainability 4. Optimized for CPU-only environments 5. Inherent explainability (glass-box model) 6. Faster computation and low latency 7. Robustness against unbalanced classes 8. Economic viability 9. Generic, expandable, and extendable
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。