arXiv:2506.03682cs.CVcs.AI2025-06被引 2

提出无需网格的图像部件相对位置学习方法,提升空间理解能力

How PARTs assemble into wholes: Learning the relative composition of images

  • 用连续相对变换建模图像部件间关系,摆脱固定网格限制
  • 在目标检测等任务中优于MAE、DropPos等网格方法
  • 适合需要精细空间感知的场景,如医学影像与视频分析

物体及其部件的组合方式,以及物体间的空间位置关系,为表征学习提供了丰富信息。因此,空间感知的自监督预训练任务受到广泛关注。现有方法多基于固定网格结构,通过预测补丁在网格中的绝对位置作为预训练目标。然而,网格方法难以捕捉真实世界中物体组合的流动性和连续性。本文提出PART,一种基于非网格补丁间连续相对变换的自监督学习方法,旨在学习图像的相对组成——即不依赖绝对外观、在部分遮挡或风格变化下仍保持一致的非网格结构化相对定位。在需要精确空间理解的任务(如目标检测和时序预测)中,PART优于基于网格的方法(如MAE和DropPos),同时在全局分类任务上保持竞争力。突破网格限制后,PART为跨数据类型(从图像到脑电图信号)的通用自监督预训练开辟新路径,在医学影像、视频和音频等领域具有应用潜力。

原文摘要 · Abstract (English)

The composition of objects and their parts, along with object-object positional relationships, provides a rich source of information for representation learning. Hence, spatial-aware pretext tasks have been actively explored in self-supervised learning. Existing works commonly start from a grid structure, where the goal of the pretext task involves predicting the absolute position index of patches within a fixed grid. However, grid-based approaches fall short of capturing the fluid and continuous nature of real-world object compositions. We introduce PART, a self-supervised learning approach that leverages continuous relative transformations between off-grid patches to overcome these limitations. By modeling how parts relate to each other in a continuous space, PART learns the relative composition of images-an off-grid structural relative positioning that is less tied to absolute appearance and can remain coherent under variations such as partial visibility or stylistic changes. In tasks requiring precise spatial understanding such as object detection and time series prediction, PART outperforms grid-based methods like MAE and DropPos, while maintaining competitive performance on global classification tasks. By breaking free from grid constraints, PART opens up a new trajectory for universal self-supervised pretraining across diverse datatypes-from images to EEG signals-with potential in medical imaging, video, and audio.

自监督学习空间建模图像理解相对位置

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。