arXiv:2509.20484cs.CV2025-09

用流式高置信度选图,少传数据也能训出好模型。

Data-Efficient Stream-Based Active Distillation for Scalable Edge Model Deployment

  • 基于置信度和多样性筛选图像,动态选最有价值样本。
  • 相同训练轮数下,模型性能接近全量数据训练。
  • 适合边缘设备实时更新,降低传输成本。

基于边缘摄像头的系统持续扩展,面临不断变化的环境,需定期更新模型。实践中,复杂的教师模型在中心服务器上对数据进行标注,再用于训练计算能力受限的边缘设备小型模型。本文研究如何选择最有效的图像进行训练,在保持低传输成本的同时最大化模型质量。结果表明,在相似训练负载(即迭代次数)下,结合高置信度流式策略与多样性方法,可实现高质量模型,且仅需极少的数据集查询。

原文摘要 · Abstract (English)

Edge camera-based systems are continuously expanding, facing ever-evolving environments that require regular model updates. In practice, complex teacher models are run on a central server to annotate data, which is then used to train smaller models tailored to the edge devices with limited computational power. This work explores how to select the most useful images for training to maximize model quality while keeping transmission costs low. Our work shows that, for a similar training load (i.e., iterations), a high-confidence stream-based strategy coupled with a diversity-based approach produces a high-quality model with minimal dataset queries.

边缘计算主动学习模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。