arXiv:2503.07637cs.LG2025-03

让解码器也用预训练模型,提升图像密集预测性能

Is Pre-training Applicable to the Decoder for Dense Prediction?

  • 设计新架构使解码器直接复用预训练权重
  • 在单目深度估计等任务上达到顶尖水平
  • 无需特殊结构,适合通用视觉任务研究者

预训练编码器广泛用于密集预测任务,因其能有效提取图像视觉特征,而解码器通常从零开始训练。由于结构差异和输入数据不同,解码器难以受益于视觉基准(如图像分类)的预训练表示。本文提出×Net,通过三项创新设计实现“预训练编码器×预训练解码器”的协作,使解码器直接利用预训练模型中的知识,融入解码过程以提升性能。仅通过耦合预训练编码器与解码器,×Net无需特定解码结构或任务定制算法,便在单目深度估计和语义分割等任务中超越先进方法,尤其在单目深度估计上取得当前最优结果。

原文摘要 · Abstract (English)

Pre-trained encoders are widely employed in dense prediction tasks for their capability to effectively extract visual features from images. The decoder subsequently processes these features to generate pixel-level predictions. However, due to structural differences and variations in input data, only encoders benefit from pre-learned representations from vision benchmarks such as image classification and self-supervised learning, while decoders are typically trained from scratch. In this paper, we introduce $\times$Net, which facilitates a "pre-trained encoder $\times$ pre-trained decoder" collaboration through three innovative designs. $\times$Net enables the direct utilization of pre-trained models within the decoder, integrating pre-learned representations into the decoding process to enhance performance in dense prediction tasks. By simply coupling the pre-trained encoder and pre-trained decoder, $\times$Net distinguishes itself as a highly promising approach. Remarkably, it achieves this without relying on decoding-specific structures or task-specific algorithms. Despite its streamlined design, $\times$Net outperforms advanced methods in tasks such as monocular depth estimation and semantic segmentation, achieving state-of-the-art performance particularly in monocular depth estimation. and semantic segmentation, achieving state-of-the-art results, especially in monocular depth estimation. embedding algorithms. Despite its streamlined design, $\times$Net outperforms advanced methods in tasks such as monocular depth estimation and semantic segmentation, achieving state-of-the-art performance particularly in monocular depth estimation.

密集预测预训练解码器深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。