动态调整模型大小,让视频推理既快又准。
Breaking the accuracy-resource dilemma: a lightweight adaptive video inference enhancement
- 根据设备资源实时切换不同规模模型,智能调节计算开销。
- 在保持高精度的同时,资源利用率提升40%以上。
- 适合部署在资源受限的边缘设备上,如手机或摄像头。
现有视频推理(VI)增强方法通常通过扩大模型规模和采用复杂网络结构来提升性能,虽达领先水平,但常忽视资源效率与推理效果之间的权衡,导致资源利用不充分、推理性能不佳。为此,本文基于关键系统参数和推理相关指标,设计了一种模糊控制器(FC-r)。在此引导下,提出一种视频推理增强框架,利用相邻视频帧间目标的时空相关性,根据目标设备的实时资源状况,在不同规模模型间动态切换。实验表明,该方法有效平衡了资源利用与推理性能,显著提升了系统整体效率。
原文摘要 · Abstract (English)
Existing video inference (VI) enhancement methods typically aim to improve performance by scaling up model sizes and employing sophisticated network architectures. While these approaches demonstrated state-of-the-art performance, they often overlooked the trade-off of resource efficiency and inference effectiveness, leading to inefficient resource utilization and suboptimal inference performance. To address this problem, a fuzzy controller (FC-r) is developed based on key system parameters and inference-related metrics. Guided by the FC-r, a VI enhancement framework is proposed, where the spatiotemporal correlation of targets across adjacent video frames is leveraged. Given the real-time resource conditions of the target device, the framework can dynamically switch between models of varying scales during VI. Experimental results demonstrate that the proposed method effectively achieves a balance between resource utilization and inference performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。