arXiv:2505.04758cs.CV2025-05中稿 · TIP 2025被引 22

轻量级模型在精度与速度间取得平衡,提升RGB-D显著目标检测性能。

Lightweight RGB-D Salient Object Detection from a Speed-Accuracy Tradeoff Perspective

  • 从深度质量、模态融合、特征表示三方面设计轻量化网络
  • 参数仅520万,推理速度达415帧/秒,超越重型模型
  • 适合移动端或实时系统部署,兼顾精度与效率

现有RGB-D方法多采用大规模骨干网络以提升精度,但牺牲效率;而现有轻量方法难以达到高精度。为平衡效率与性能,本文提出速度-精度权衡网络SATNet,从深度质量、模态融合和特征表示三方面入手。针对深度质量,引入Depth Anything Model生成高质量深度图,缓解当前数据集中的多模态差异。在模态融合方面,提出解耦注意力模块(DAM),将多模态特征解耦为双视角向量,提取判别性信息。在特征表示上,设计双向逆向框架的双重信息表示模块(DIRM),建模纹理与显著性特征,通过双向反向传播优化参数。最后在解码器中设计双重特征聚合模块(DFAM)融合纹理与显著性特征。在五个公开的RGB-D SOD数据集上的大量实验表明,SATNet优于当前SOTA的基于CNN的重型模型,同时实现仅5.2M参数和415 FPS的轻量级架构。

原文摘要 · Abstract (English)

Current RGB-D methods usually leverage large-scale backbones to improve accuracy but sacrifice efficiency. Meanwhile, several existing lightweight methods are difficult to achieve high-precision performance. To balance the efficiency and performance, we propose a Speed-Accuracy Tradeoff Network (SATNet) for Lightweight RGB-D SOD from three fundamental perspectives: depth quality, modality fusion, and feature representation. Concerning depth quality, we introduce the Depth Anything Model to generate high-quality depth maps,which effectively alleviates the multi-modal gaps in the current datasets. For modality fusion, we propose a Decoupled Attention Module (DAM) to explore the consistency within and between modalities. Here, the multi-modal features are decoupled into dual-view feature vectors to project discriminable information of feature maps. For feature representation, we develop a Dual Information Representation Module (DIRM) with a bi-directional inverted framework to enlarge the limited feature space generated by the lightweight backbones. DIRM models texture features and saliency features to enrich feature space, and employ two-way prediction heads to optimal its parameters through a bi-directional backpropagation. Finally, we design a Dual Feature Aggregation Module (DFAM) in the decoder to aggregate texture and saliency features. Extensive experiments on five public RGB-D SOD datasets indicate that the proposed SATNet excels state-of-the-art (SOTA) CNN-based heavyweight models and achieves a lightweight framework with 5.2 M parameters and 415 FPS.

RGB-D轻量化显著目标检测高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。