arXiv:2506.07806cs.LGcs.CV2025-06ICML

在视角模糊下实现可识别的物体表征,无需视角标注

Identifiable Object Representations under Spatial Ambiguities

  • 多视角概率方法聚合视图槽,学习不变内容与视角信息
  • 理论保证可识别性,在标准与新数据集上表现稳健
  • 适合需要鲁棒物体理解的视觉系统开发

模块化、以物体为中心的表征对类人推理至关重要,但在遮挡和视角模糊等空间不确定性下难以获取。现有方法面临理论与实践双重挑战。本文提出一种新型多视角概率方法,通过聚合各视角特定的槽位,捕捉不变内容信息,同时学习解耦的全局视角级信息。相比以往单视角方法,本方法能有效解决空间模糊问题,提供可识别性的理论保障,且无需视角标注。在标准基准和新构建的复杂数据集上的大量实验验证了该方法的鲁棒性与可扩展性。

原文摘要 · Abstract (English)

Modular object-centric representations are essential for *human-like reasoning* but are challenging to obtain under spatial ambiguities, *e.g. due to occlusions and view ambiguities*. However, addressing challenges presents both theoretical and practical difficulties. We introduce a novel multi-view probabilistic approach that aggregates view-specific slots to capture *invariant content* information while simultaneously learning disentangled global *viewpoint-level* information. Unlike prior single-view methods, our approach resolves spatial ambiguities, provides theoretical guarantees for identifiability, and requires *no viewpoint annotations*. Extensive experiments on standard benchmarks and novel complex datasets validate our method's robustness and scalability.

物体表征多视角可识别性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。