CLIDD通过分层独立采样实现高效高精度局部特征描述,适合实时空间智能任务。
CLIDD: Cross-Layer Independent Deformable Description for Efficient and Discriminative Local Feature Representation
- 分层独立采样结合可学习偏移,捕捉多尺度细粒结构。
- 超小模型仅用0.004M参数,精度媲美SuperPoint,模型缩小99.7%。
- 支持边缘设备实时运行(>200 FPS),适配不同部署需求。
鲁棒的局部特征表示对机器人导航、增强现实等空间智能任务至关重要。可靠的对应关系需要兼具高区分度与计算效率的描述子。为此,我们提出跨层独立可变形描述(CLIDD),通过直接从独立特征层级采样实现卓越的区分性。该方法利用可学习偏移捕捉多尺度细粒度结构细节,同时避免统一密集表示带来的计算负担。为保证实时性能,采用硬件感知的核融合策略以最大化推理吞吐量。此外,构建了可扩展框架,结合轻量级架构与融合度量学习和知识蒸馏的训练协议,生成一系列针对不同部署约束优化的模型变体。大量实验表明,本方法在匹配精度与计算效率上均表现优异:超紧凑变体仅需0.004M参数,精度媲美SuperPoint,模型尺寸减少99.7%;高性能配置超越所有现有SOTA方法(包括高容量DINOv2基框架),并在边缘设备上实现超过200 FPS的推理速度。结果证明,CLIDD在极低计算开销下实现了高精度局部特征匹配,为实时空间智能任务提供稳健可扩展的解决方案。
原文摘要 · Abstract (English)
Robust local feature representations are essential for spatial intelligence tasks such as robot navigation and augmented reality. Establishing reliable correspondences requires descriptors that provide both high discriminative power and computational efficiency. To address this, we introduce Cross-Layer Independent Deformable Description (CLIDD), a method that achieves superior distinctiveness by sampling directly from independent feature hierarchies. This approach utilizes learnable offsets to capture fine-grained structural details across scales while bypassing the computational burden of unified dense representations. To ensure real-time performance, we implement a hardware-aware kernel fusion strategy that maximizes inference throughput. Furthermore, we develop a scalable framework that integrates lightweight architectures with a training protocol leveraging both metric learning and knowledge distillation. This scheme generates a wide spectrum of model variants optimized for diverse deployment constraints. Extensive evaluations demonstrate that our approach achieves superior matching accuracy and exceptional computational efficiency simultaneously. Specifically, the ultra-compact variant matches the precision of SuperPoint while utilizing only 0.004M parameters, achieving a 99.7% reduction in model size. Furthermore, our high-performance configuration outperforms all current state-of-the-art methods, including high-capacity DINOv2-based frameworks, while exceeding 200 FPS on edge devices. These results demonstrate that CLIDD delivers high-precision local feature matching with minimal computational overhead, providing a robust and scalable solution for real-time spatial intelligence tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。