arXiv:2504.08912cs.LGcs.AI2025-04被引 13

构建首个通用双曲基础模型框架,支持多模态高效建模。

HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules

  • 提供可组合的双曲核心模块,免去从零重构
  • 首例全双曲视觉Transformer性能超越欧氏模型
  • 适合研究双曲神经网络与多模态表征的学者

双曲神经网络在建模分层数据方面表现出强大能力。近期研究表明,基础模型中标记分布具有无标度特性,表明双曲空间比欧几里得空间更适合多数预训练与下游任务。然而,现有工具缺乏构建双曲基础模型的核心组件,难以利用最新进展。我们提出 HyperCore,一个开源的综合性框架,提供跨多模态构建双曲基础模型的核心模块。其模块可轻松组合,用于开发新型双曲基础模型,无需从头修改欧氏模块,避免重复研究。为验证其泛化性,我们构建并测试了首个全双曲视觉变换器(LViT)和微调流程、首个全双曲多模态CLIP模型(L-CLIP),以及基于双曲图编码器的混合图RAG。实验表明,LViT性能优于其欧氏对应模型。此外,我们在双曲GNN、CNN、Transformer和视觉变换器上进行基准测试与复现,凸显HyperCore的优势。

原文摘要 · Abstract (English)

Hyperbolic neural networks have emerged as a powerful tool for modeling hierarchical data across diverse modalities. Recent studies show that token distributions in foundation models exhibit scale-free properties, suggesting that hyperbolic space is a more suitable ambient space than Euclidean space for many pre-training and downstream tasks. However, existing tools lack essential components for building hyperbolic foundation models, making it difficult to leverage recent advancements. We introduce HyperCore, a comprehensive open-source framework that provides core modules for constructing hyperbolic foundation models across multiple modalities. HyperCore's modules can be effortlessly combined to develop novel hyperbolic foundation models, eliminating the need to extensively modify Euclidean modules from scratch and possible redundant research efforts. To demonstrate its versatility, we build and test the first fully hyperbolic vision transformers (LViT) with a fine-tuning pipeline, the first fully hyperbolic multimodal CLIP model (L-CLIP), and a hybrid Graph RAG with a hyperbolic graph encoder. Our experiments demonstrate that LViT outperforms its Euclidean counterpart. Additionally, we benchmark and reproduce experiments across hyperbolic GNNs, CNNs, Transformers, and vision Transformers to highlight HyperCore's advantages.

双曲神经网络基础模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。