arXiv:2602.05330cs.CV2026-02International Conf…被引 5

无需标注数据,用透视模型生成全景图的多任务理解新方法。

MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction Priors

  • 用透视模型生成伪标签,解决全景图标注稀缺问题。
  • 分组处理旋转不变与变化任务,提升多任务协同效果。
  • 适合做全景视觉理解、跨任务学习的研究者参考。

全景场景理解对沉浸式应用至关重要,但因高分辨率多任务标注稀缺而困难重重。尽管透视基础模型通过数据规模取得成功,直接适配全景域常因严重几何失真和坐标系统差异失败。此外,球面空间中各类密集预测任务间的关联尚不明确。为此,我们提出MTPano,一种通过无标签训练流程构建的鲁棒多任务全景基础模型。首先,为克服数据稀缺,我们利用强大的透视密集先验:将全景图像投影至透视块,使用现成的基础模型生成准确、域间隙无差的伪标签,并重新投影作为块级监督。其次,为缓解任务间干扰,我们将任务分为旋转不变(如深度、分割)与旋转相关(如表面法线)两类,引入全景双路桥网络,通过注入绝对位置与射线方向先验的几何感知调制层解耦特征流。为应对等距矩形投影(ERP)的畸变,我们结合ERP token mixer与双分支桥网络,配合梯度截断机制,促进有益的跨任务信息共享,同时阻断不兼容任务属性的冲突梯度。此外,引入辅助任务以增强跨任务学习。大量实验表明,MTPano在多个基准上达到领先性能,且优于特定任务的全景专用基础模型。

原文摘要 · Abstract (English)

Comprehensive panoramic scene understanding is critical for immersive applications, yet it remains challenging due to the scarcity of high-resolution, multi-task annotations. While perspective foundation models have achieved success through data scaling, directly adapting them to the panoramic domain often fails due to severe geometric distortions and coordinate system discrepancies. Furthermore, the underlying relations between diverse dense prediction tasks in spherical spaces are underexplored. To address these challenges, we propose MTPano, a robust multi-task panoramic foundation model established by a label-free training pipeline. First, to circumvent data scarcity, we leverage powerful perspective dense priors. We project panoramic images into perspective patches to generate accurate, domain-gap-free pseudo-labels using off-the-shelf foundation models, which are then re-projected to serve as patch-wise supervision. Second, to tackle the interference between task types, we categorize tasks into rotation-invariant (e.g., depth, segmentation) and rotation-variant (e.g., surface normals) groups. We introduce the Panoramic Dual BridgeNet, which disentangles these feature streams via geometry-aware modulation layers that inject absolute position and ray direction priors. To handle the distortion from equirectangular projections (ERP), we incorporate ERP token mixers followed by a dual-branch BridgeNet for interactions with gradient truncation, facilitating beneficial cross-task information sharing while blocking conflicting gradients from incompatible task attributes. Additionally, we introduce auxiliary tasks to fertilize the cross-task learning process. Extensive experiments demonstrate that MTPano achieves state-of-the-art performance on multiple benchmarks and delivers competitive results against task-specific panoramic specialist foundation models.

全景理解多任务学习伪标签扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。