arXiv:2512.18192cs.CV2025-12

用显式图结构提升多部件物体识别的鲁棒性。

Multi-Part Object Representations via Graph Structures and Co-Part Discovery

  • 构建显式图结构表示物体部件,明确部件间关系。
  • 在遮挡和分布外场景下,物体识别准确率显著提升。
  • 适合需要强泛化能力的视觉模型研究者使用。

从图像中发现以物体为中心的表征可显著提升视觉模型的鲁棒性、样本效率和泛化能力。现有针对多部件物体的方法通常采用隐式表征,无法在遮挡或分布外场景中有效识别所学物体,原因在于其假设部件-整体关系通过间接训练目标隐含编码于表征中。本文提出一种新方法,利用显式图结构表示部件,并设计协同部件发现算法。同时引入三个基准测试,评估物体中心方法在遮挡和分布外场景下的鲁棒性。在模拟、真实及现实图像上的实验结果表明,本方法在发现物体质量上优于当前最优方法,并能准确识别遮挡与分布外场景中的多部件物体。此外,所发现的物体中心表征在下游任务中对关键物体属性预测更准确,展示了该方法推动物体中心表征领域发展的潜力。

原文摘要 · Abstract (English)

Discovering object-centric representations from images can significantly enhance the robustness, sample efficiency and generalizability of vision models. Works on images with multi-part objects typically follow an implicit object representation approach, which fail to recognize these learned objects in occluded or out-of-distribution contexts. This is due to the assumption that object part-whole relations are implicitly encoded into the representations through indirect training objectives. We address this limitation by proposing a novel method that leverages on explicit graph representations for parts and present a co-part object discovery algorithm. We then introduce three benchmarks to evaluate the robustness of object-centric methods in recognizing multi-part objects within occluded and out-of-distribution settings. Experimental results on simulated, realistic, and real-world images show marked improvements in the quality of discovered objects compared to state-of-the-art methods, as well as the accurate recognition of multi-part objects in occluded and out-of-distribution contexts. We also show that the discovered object-centric representations can more accurately predict key object properties in a downstream task, highlighting the potential of our method to advance the field of object-centric representations.

物体表征图神经网络多部件识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。