arXiv:2510.07089cs.CV2025-10被引 1

用深度注意力机制自动发现图像中的物体,无需人工标注

DADO: A Depth-Attention framework for Object Discovery

  • 结合注意力与深度信息,动态调整特征权重
  • 在标准数据集上精度超越现有方法,无需微调
  • 适合无监督目标发现任务的研究与应用

无监督对象发现旨在不依赖人工标注的情况下识别和定位图像中的对象,仍是计算机视觉中的重大挑战。本文提出一种新模型 DADO(Depth-Attention self-supervised technique for Discovering unseen Objects),融合注意力机制与深度模型,以识别图像中潜在的对象。为应对注意力图噪声或包含多深度层级的复杂场景问题,DADO 采用动态加权策略,根据每张图像的全局特征自适应地强调注意力或深度特征。我们在标准基准上评估了 DADO,结果表明其在对象发现准确率和鲁棒性方面优于现有最先进方法,且无需微调。

原文摘要 · Abstract (English)

Unsupervised object discovery, the task of identifying and localizing objects in images without human-annotated labels, remains a significant challenge and a growing focus in computer vision. In this work, we introduce a novel model, DADO (Depth-Attention self-supervised technique for Discovering unseen Objects), which combines an attention mechanism and a depth model to identify potential objects in images. To address challenges such as noisy attention maps or complex scenes with varying depth planes, DADO employs dynamic weighting to adaptively emphasize attention or depth features based on the global characteristics of each image. We evaluated DADO on standard benchmarks, where it outperforms state-of-the-art methods in object discovery accuracy and robustness without the need for fine-tuning.

无监督学习目标发现注意力机制深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。