arXiv:2607.25531cs.LGcs.AI2026-07

提出多尺度结构特征,实现可解释的持续视觉识别。

Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework

论文配图:Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework
图 1 · 摘自论文原文
  • 用多尺度结构特征捕捉形状边缘与空间关系
  • 准确率超越基线,且不存储历史数据
  • 适合需要可解释性与持续学习的应用场景

当前机器学习难以持续学习、复用旧知识并展现可理解的内部结构。一种新型发展式无梯度学习框架通过局部变异与选择构建输入的离散拓扑模型,天然保证持续学习:新观测仅优化现有结构而不覆盖旧知识,无需回放缓冲或预定义任务边界。其视觉扩展在形状识别中验证了该原理,但依赖表达力有限的特征表示,限制了识别准确率。本文引入一种新的多尺度视觉特征表示,编码跨尺度的形状结构,包含边缘、轮廓及其空间关系,并整合至网络优化学习过程;同时改进学习动态与预测读出机制。研究聚焦二维形状,以类增量MNIST为可控可解释基准,直接测量持续学习行为。所提方法显著提升准确率,达到或超过基于重放与正则化的基线,在相似存储开销下不存储任何历史数据,且保持框架核心特性:早期学习类别在引入新类时仍被保留,无破坏性适应,学习表征仍可人工解读。差异在于保留能力:基线在训练周期内丢失大部分刚学类别,需后续重学,而本方法避免此问题。意义在于学习方式:系统逐样本整合信息,同时严格保持对已有样本的响应。

原文摘要 · Abstract (English)

Contemporary machine learning struggles to learn continually, reuse prior knowledge, and expose a comprehensible internal structure. A recently proposed developmental, gradient-free learning framework addresses these limitations by learning a discrete, topological model of its inputs through local variation and selection, yielding an inherent continual-learning guarantee: new observations refine existing structure without overwriting past knowledge, and without replay buffers or predefined task boundaries. Its extension to visual inputs demonstrated this principle on shape recognition, but relied on a feature representation of limited expressivity that capped recognition accuracy. We introduce a new visual feature representation that encodes shape structure across multiple scales, capturing edge and contour features together with their spatial relations, and integrate it with the network-refinement learning process; we further improve the learning dynamics and the read-out used to predict from the learned model. The study targets two-dimensional shape, with class-incremental MNIST as a controlled, interpretable benchmark in which continual-learning behavior can be measured directly. Our approach substantially increases accuracy over the prior representation, matching or exceeding replay- and regularisation-based baselines at comparable storage while storing no past data, and preserves the framework's defining behavior: earlier-learned classes are retained as new ones are introduced, with no destructive adaptation, and the learned representations remain human-interpretable. What separates the methods is retention: the baselines surrender most of a just-trained class within its own cycle and relearn it afterwards, which ours does not. The significance lies in the manner of learning. The system integrates information one sample at a time while provably preserving its responses to...

持续学习可解释性多尺度特征视觉识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。