arXiv:2412.15150cs.CVcs.AI2024-12

用颜色空间改进无监督目标检测,提升图像重建与特征解耦效果

Leveraging Color Channel Independence for Improved Unsupervised Object Detection

  • 将预测目标转为RGB-S色空,融合饱和度信息增强表征能力
  • 在5个数据集上实现更优的图像重建与特征解耦性能
  • 方法零开销、通用性强,适用于多种视觉任务与模型架构

以对象为中心的架构可从视觉场景中学习独立的对象表征,支持下游对象级应用。类似基于自编码器的图像模型,当前对象中心方法通常在由RGB色彩空间编码的图像上使用无监督重建损失进行训练。本文挑战了‘RGB图像是计算机视觉无监督学习最优色彩空间’这一常见假设。我们从概念和实证角度指出,如HSV等其他色彩空间具备对光照条件更鲁棒等关键特性,有利于对象中心表征学习。进一步发现,要求模型预测额外颜色通道可显著提升性能。为此,我们提出将预测目标转换至RGB-S空间——在RGB基础上加入HSV的饱和度分量——在五个常见评估数据集上实现了明显更好的重建效果与特征解耦能力。该复合色彩空间方法几乎无计算开销,与模型架构无关,且广泛适用于各类视觉计算任务与训练方式。本方法的发现为超越对象中心学习的计算机视觉任务研究提供了新思路。

原文摘要 · Abstract (English)

Object-centric architectures can learn to extract distinct object representations from visual scenes, enabling downstream applications on the object level. Similarly to autoencoder-based image models, object-centric approaches have been trained on the unsupervised reconstruction loss of images encoded by RGB color spaces. In our work, we challenge the common assumption that RGB images are the optimal color space for unsupervised learning in computer vision. We discuss conceptually and empirically that other color spaces, such as HSV, bear essential characteristics for object-centric representation learning, like robustness to lighting conditions. We further show that models improve when requiring them to predict additional color channels. Specifically, we propose to transform the predicted targets to the RGB-S space, which extends RGB with HSV's saturation component and leads to markedly better reconstruction and disentanglement for five common evaluation datasets. The use of composite color spaces can be implemented with basically no computational overhead, is agnostic of the models' architecture, and is universally applicable across a wide range of visual computing tasks and training types. The findings of our approach encourage additional investigations in computer vision tasks beyond object-centric learning.

无监督学习对象检测色彩空间特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。