arXiv:2508.19864cs.CV2025-08

通过多尺度结构化表示学习,提升自监督视觉模型对物体的感知能力。

Self-supervised structured object representation learning

  • 基于新模块ProtoScale,分步实现语义分组、实例分离和层级构建
  • 在COCO与UA-DETRAC数据集上,少标注数据下检测性能超越当前最佳
  • 保留完整场景上下文,更适合密集预测任务,适合目标检测研究者

自监督学习(SSL)已成为视觉表征学习的重要方法。尽管现有方法在全局图像理解上表现优异,但在捕捉场景中结构化表示方面仍存在局限。本文提出一种新型自监督方法,通过逐步融合语义分组、实例级分离与层级结构构建,实现对视觉结构的渐进式建模。该方法基于创新的ProtoScale模块,能捕捉多空间尺度下的视觉元素。不同于依赖随机裁剪与全局嵌入的DINO等策略,本方法在增强视图中保留完整场景上下文,显著提升密集预测任务性能。我们在结合多个数据集(COCO与UA-DETRAC)的子集上验证方法有效性。实验结果表明,所学表示具有以物体为中心的特性,能有效增强监督检测性能,在标注数据有限及微调轮次较少的情况下,仍优于现有最优方法。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has emerged as a powerful technique for learning visual representations. While recent SSL approaches achieve strong results in global image understanding, they are limited in capturing the structured representation in scenes. In this work, we propose a self-supervised approach that progressively builds structured visual representations by combining semantic grouping, instance level separation, and hierarchical structuring. Our approach, based on a novel ProtoScale module, captures visual elements across multiple spatial scales. Unlike common strategies like DINO that rely on random cropping and global embeddings, we preserve full scene context across augmented views to improve performance in dense prediction tasks. We validate our method on downstream object detection tasks using a combined subset of multiple datasets (COCO and UA-DETRAC). Experimental results show that our method learns object centric representations that enhance supervised object detection and outperform the state-of-the-art methods, even when trained with limited annotated data and fewer fine-tuning epochs.

自监督学习物体检测结构化表征多尺度建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。