arXiv:2603.22387cs.CV2026-03被引 6

小模型也能全任务通用,高效感知编码器让边缘设备智能更强大

Efficient Universal Perception Encoder

  • 先放大再压缩:用统一大教师模型提炼多领域专家知识
  • 同等规模下性能超越单领域专家,且优于以往集成方法
  • 适合部署在算力有限的手机、IoT等边缘设备上

在智能边缘设备上运行AI可带来多样化的用户体验,但受限于计算资源且需同时处理多种任务,亟需小型化却具备强大通用表征能力的视觉编码器。本文提出高效通用感知编码器(EUPE),通过从多个领域专家基础视觉编码器中进行知识蒸馏实现。与以往直接从多个教师模型合并压缩的方法不同,我们证明了先将多教师知识整合为一个大型代理教师,再从中蒸馏出小型编码器更为有效。实验表明,EUPE在同等模型尺寸下,在多个下游任务域上表现不逊于甚至优于独立的领域专家,并显著优于以往的聚合式蒸馏方法。我们已公开EUPE完整模型系列及代码,以促进后续研究。

原文摘要 · Abstract (English)

Running AI models on smart edge devices can unlock versatile user experiences, but presents challenges due to limited compute and the need to handle multiple tasks simultaneously. This requires a vision encoder with small size but powerful and versatile representations. We present our method, Efficient Universal Perception Encoder (EUPE), which offers both inference efficiency and universally good representations for diverse downstream tasks. We achieve this by distilling from multiple domain-expert foundation vision encoders. Unlike previous agglomerative methods that directly scale down from multiple teachers to an efficient encoder, we demonstrate the importance of first scaling up to a large proxy teacher and then scaling down from this single teacher. Experiments show that EUPE achieves on-par or better performance than individual domain experts of the same size on diverse task domains and also outperforms previous agglomerative encoders. We release the full family of EUPE models and the code to foster future research.

视觉编码器边缘计算知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。