模仿生物遗传机制,让小模型高效继承大模型的有用知识。
HKT: A Biologically Inspired Framework for Modular Hereditary Knowledge Transfer in Neural Networks
- 用三阶段生物启发框架提取、传递、融合关键特征
- 在多个视觉任务中提升小模型性能,超越传统蒸馏方法
- 适合资源受限场景下的高性能模型部署
当前神经网络研究倾向于通过增加深度和容量来提升性能,但往往牺牲可集成性与效率。本文提出一种优化小型可部署模型的方法,通过结构化知识继承增强其能力。我们引入仿生框架赫氏知识迁移(HKT),实现从大型预训练父模型到小型子模型的模块化、选择性任务相关特征迁移。不同于标准知识蒸馏中对教师输出的均匀模仿,HKT借鉴扁虫记忆RNA传递等生物遗传机制,驱动多阶段特征迁移过程。将神经网络模块视为功能载体,知识通过提取、传递、混合(ETM)三个生物启发组件完成传输。一种新型遗传注意力(GA)机制调控继承与原生表征的融合,确保对齐与选择性。我们在多种视觉任务上评估HKT,包括光流(Sintel、KITTI)、图像分类(CIFAR-10)和语义分割(LiTS),结果表明其显著提升子模型性能的同时保持紧凑性,且持续优于传统蒸馏方法,提供了一种通用、可解释、可扩展的高绩效神经网络部署方案。
原文摘要 · Abstract (English)
A prevailing trend in neural network research suggests that model performance improves with increasing depth and capacity - often at the cost of integrability and efficiency. In this paper, we propose a strategy to optimize small, deployable models by enhancing their capabilities through structured knowledge inheritance. We introduce Hereditary Knowledge Transfer (HKT), a biologically inspired framework for modular and selective transfer of task-relevant features from a larger, pretrained parent network to a smaller child model. Unlike standard knowledge distillation, which enforces uniform imitation of teacher outputs, HKT draws inspiration from biological inheritance mechanisms - such as memory RNA transfer in planarians - to guide a multi-stage process of feature transfer. Neural network blocks are treated as functional carriers, and knowledge is transmitted through three biologically motivated components: Extraction, Transfer, and Mixture (ETM). A novel Genetic Attention (GA) mechanism governs the integration of inherited and native representations, ensuring both alignment and selectivity. We evaluate HKT across diverse vision tasks, including optical flow (Sintel, KITTI), image classification (CIFAR-10), and semantic segmentation (LiTS), demonstrating that it significantly improves child model performance while preserving its compactness. The results show that HKT consistently outperforms conventional distillation approaches, offering a general-purpose, interpretable, and scalable solution for deploying high-performance neural networks in resource-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。