通过跨视图相关性提升多任务学习的3D感知能力,增强场景理解。
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
- 引入跨视图模块捕捉多视角几何一致性,增强3D感知。
- 在NYUv2和PASCAL-Context上显著提升分割与深度估计性能。
- 模块轻量且通用,适用于单/多视图数据,可集成到现有网络中。
本文针对单个网络联合执行多种密集预测任务(如分割与深度估计)的多任务学习(MTL)挑战。现有方法主要在2D图像空间建模跨任务关联,常导致特征缺乏3D感知。我们提出,3D感知对建立全面场景理解所必需的跨任务关联至关重要。为此,我们在MTL网络中引入跨视图相关性(即代价体)作为几何一致性先验。具体地,设计一个轻量级跨视图模块(CvM),在任务间共享,用于跨视图信息交换并捕捉跨视图相关性,与MTL编码器提取的特征融合以实现多任务预测。该模块架构无关,适用于单视图与多视图数据。在NYUv2和PASCAL-Context上的大量实验表明,该方法有效将几何一致性注入现有MTL方法,显著提升性能。
原文摘要 · Abstract (English)
This paper addresses the challenge of training a single network to jointly perform multiple dense prediction tasks, such as segmentation and depth estimation, i.e., multi-task learning (MTL). Current approaches mainly capture cross-task relations in the 2D image space, often leading to unstructured features lacking 3D-awareness. We argue that 3D-awareness is vital for modeling cross-task correlations essential for comprehensive scene understanding. We propose to address this problem by integrating correlations across views, i.e., cost volume, as geometric consistency in the MTL network. Specifically, we introduce a lightweight Cross-view Module (CvM), shared across tasks, to exchange information across views and capture cross-view correlations, integrated with a feature from MTL encoder for multi-task predictions. This module is architecture-agnostic and can be applied to both single and multi-view data. Extensive results on NYUv2 and PASCAL-Context demonstrate that our method effectively injects geometric consistency into existing MTL methods to improve performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。