arXiv:2506.09895cs.CV2025-06ICCV

不用预测器的胶囊网络,让模型自动学会物体姿态变化规律。

EquiCaps: Predictor-Free Pose-Aware Pre-Trained Capsule Networks

  • 用胶囊网络内在的姿态感知能力替代专用预测器
  • 在3D旋转预测上达到0.78的R²,超越现有方法0.04以上
  • 适合研究姿态不变性、自监督学习的学者参考

学习对变换保持不变且等变的自监督表征是推动视觉任务超越传统分类的关键。然而,许多方法依赖专用预测器来编码等变性,而胶囊网络因其天然具备可解释的姿态感知能力,展现出更优潜力。为此,本文提出EquiCaps(等变胶囊网络),一种基于胶囊的无预测器姿态感知自监督方法,利用胶囊自身的姿态敏感特性提升姿态估计性能。为进一步验证假设,引入3DIEBench-T,扩展了3D物体渲染基准数据集,支持多几何变换下的复杂任务评估。实验表明,EquiCaps在旋转预测任务上达到0.78的监督级R²,较SIE和CapsIE分别提升0.05和0.04。相比非胶囊类等变方法,EquiCaps在复合几何变换下仍保持稳健表现,凸显其泛化能力及无预测器胶囊架构的前景。

原文摘要 · Abstract (English)

Learning self-supervised representations that are invariant and equivariant to transformations is crucial for advancing beyond traditional visual classification tasks. However, many methods rely on predictor architectures to encode equivariance, despite evidence that architectural choices, such as capsule networks, inherently excel at learning interpretable pose-aware representations. To explore this, we introduce EquiCaps (Equivariant Capsule Network), a capsule-based approach to pose-aware self-supervision that eliminates the need for a specialised predictor for enforcing equivariance. Instead, we leverage the intrinsic pose-awareness capabilities of capsules to improve performance in pose estimation tasks. To further challenge our assumptions, we increase task complexity via multi-geometric transformations to enable a more thorough evaluation of invariance and equivariance by introducing 3DIEBench-T, an extension of a 3D object-rendering benchmark dataset. Empirical results demonstrate that EquiCaps outperforms prior state-of-the-art equivariant methods on rotation prediction, achieving a supervised-level $R^2$ of 0.78 on the 3DIEBench rotation prediction benchmark and improving upon SIE and CapsIE by 0.05 and 0.04 $R^2$, respectively. Moreover, in contrast to non-capsule-based equivariant approaches, EquiCaps maintains robust equivariant performance under combined geometric transformations, underscoring its generalisation capabilities and the promise of predictor-free capsule architectures.

胶囊网络自监督姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。