通过参数统计自动剪枝视觉特征,提升机器人感知速度与效率
ISCS: Parameter-Guided Feature Pruning for Resource-Constrained Embodied Perception
- 利用模型参数方差和偏置估计通道重要性,无需耗时实验
- 在多个数据集上实现低延迟推理,端到端延迟显著降低
- 适合资源受限的实时人机交互机器人系统使用
具身智能领域的研究一致表明,鲁棒感知对人机交互至关重要,但将高保真视觉模型部署在计算资源有限的设备上仍面临挑战,受限于本地算力与传输延迟。尽管潜在表示中的冗余可提升系统效率,现有方法通常依赖代价高昂的数据集特异性消融实验或不适合实时边缘-机器人协同的复杂熵模型。本文提出一种通用、数据无关的方法,用于识别并选择性传输预训练编码器中的结构关键通道。不同于暴力试错,本方法基于内在参数统计——权重方差与偏置——估算通道重要性。分析揭示出一种稳定组织结构,称为不变显著通道空间(ISCS):显著核心通道捕捉主要结构,显著辅助通道编码精细视觉细节。基于ISCS,我们设计了一种确定性静态剪枝策略,支持轻量级分片计算。跨多个数据集的实验表明,该方法避免了复杂的熵建模,实现了确定性的超低延迟流水线,显著降低端到端延迟,为资源受限的人类感知具身系统提供了关键的速度-精度权衡。
原文摘要 · Abstract (English)
Prior studies in embodied AI consistently show that robust perception is critical for human-robot interaction, yet deploying high-fidelity visual models on resource-constrained agents remains challenging due to limited on-device computation power and transmission latency. Exploiting the redundancy in latent representations could improve system efficiency, yet existing approaches often rely on costly dataset-specific ablation tests or heavy entropy models unsuitable for real-time edge-robot collaboration. We propose a generalizable, dataset-agnostic method to identify and selectively transmit structure-critical channels in pretrained encoders. Instead of brute-force empirical evaluations, our approach leverages intrinsic parameter statistics-weight variances and biases-to estimate channel importance. This analysis reveals a consistent organizational structure, termed the Invariant Salient Channel Space (ISCS), where Salient-Core channels capture dominant structures while Salient-Auxiliary channels encode fine visual details. Building on ISCS, we introduce a deterministic static pruning strategy that enables lightweight split-computing. Experiments across different datasets demonstrate that our method achieves a deterministic, ultra-low latency pipeline by bypassing heavy entropy modeling. Our method reduces end-to-end latency, providing a critical speed-accuracy trade-off for resource-constrained human-aware embodied systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。