arXiv:2409.16897cs.CV2024-09被引 11

将视觉Transformer引入双曲空间,提升层次结构建模能力。

HVT: A Comprehensive Vision Framework for Learning in Non-Euclidean Space

  • 用双曲距离和莫比乌斯变换改进自注意力机制
  • 在ImageNet上实现优于传统ViT的分类性能
  • 适合处理具有层级关系的图像数据

非欧几里得空间中的数据表示已被证明能有效捕捉现实世界数据中的层次化与复杂关系。特别是双曲空间,能高效嵌入层次结构。本文提出双曲视觉Transformer(HVT),是视觉Transformer(ViT)的新扩展,整合了双曲几何。传统ViT运行于欧氏空间,而本方法通过引入双曲距离和莫比乌斯变换,增强自注意力机制,更有效地建模图像数据中的层次与关联依赖。我们给出了严格的数学推导,展示如何将双曲几何融入注意力层、前馈网络及优化过程。在ImageNet数据集上,该方法实现了更优的图像分类性能。

原文摘要 · Abstract (English)

Data representation in non-Euclidean spaces has proven effective for capturing hierarchical and complex relationships in real-world datasets. Hyperbolic spaces, in particular, provide efficient embeddings for hierarchical structures. This paper introduces the Hyperbolic Vision Transformer (HVT), a novel extension of the Vision Transformer (ViT) that integrates hyperbolic geometry. While traditional ViTs operate in Euclidean space, our method enhances the self-attention mechanism by leveraging hyperbolic distance and Möbius transformations. This enables more effective modeling of hierarchical and relational dependencies in image data. We present rigorous mathematical formulations, showing how hyperbolic geometry can be incorporated into attention layers, feed-forward networks, and optimization. We offer improved performance for image classification using the ImageNet dataset.

视觉Transformer双曲几何图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。