arXiv:2603.07570cs.CV2026-03被引 3

多任务自适应学习+跨维度特征引导,提升RGB-D场景理解效率

Efficient RGB-D Scene Understanding via Multi-task Adaptive Learning and Cross-dimensional Feature Guidance

论文配图:Efficient RGB-D Scene Understanding via Multi-task Adaptive Learning and Cross-dimensional Feature Guidance
图 1 · 摘自论文原文
  • 设计多任务自适应损失函数,动态调整不同任务的学习策略
  • 在NYUv2等数据集上实现更高分割精度与更快处理速度
  • 适用于机器人感知、自动驾驶等需高效场景理解的场景

场景理解在机器人智能与自主性中起关键作用。传统方法常面临遮挡、边界模糊及难以根据任务需求和样本差异自适应注意力等问题。本文提出一种高效的RGB-D场景理解模型,可同时完成语义分割、实例分割、朝向估计、全景分割和场景分类。模型采用增强型融合编码器,有效利用RGB与深度输入中的冗余信息。针对语义分割,引入归一化焦点通道层和上下文特征交互层,缓解浅层特征误导和局部-全局特征表征不足问题。实例分割采用非瓶颈1D结构,在参数更少的情况下实现更优轮廓表达。此外,提出多任务自适应损失函数,根据场景变化动态调整各任务学习策略。在NYUv2、SUN RGB-D和Cityscapes数据集上的大量实验表明,该方法在分割准确率和处理速度上均优于现有方法。

原文摘要 · Abstract (English)

Scene understanding plays a critical role in enabling intelligence and autonomy in robotic systems. Traditional approaches often face challenges, including occlusions, ambiguous boundaries, and the inability to adapt attention based on task-specific requirements and sample variations. To address these limitations, this paper presents an efficient RGB-D scene understanding model that performs a range of tasks, including semantic segmentation, instance segmentation, orientation estimation, panoptic segmentation, and scene classification. The proposed model incorporates an enhanced fusion encoder, which effectively leverages redundant information from both RGB and depth inputs. For semantic segmentation, we introduce normalized focus channel layers and a context feature interaction layer, designed to mitigate issues such as shallow feature misguidance and insufficient local-global feature representation. The instance segmentation task benefits from a non-bottleneck 1D structure, which achieves superior contour representation with fewer parameters. Additionally, we propose a multi-task adaptive loss function that dynamically adjusts the learning strategy for different tasks based on scene variations. Extensive experiments on the NYUv2, SUN RGB-D, and Cityscapes datasets demonstrate that our approach outperforms existing methods in both segmentation accuracy and processing speed.

RGB-D多任务学习场景理解分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。