arXiv:2512.08820cs.CVcs.AI2025-12中稿 · IEEE Transactions …

无需训练的双曲适配器提升跨模态推理能力

Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning

  • 在双曲空间中建模视觉语言语义层级关系
  • 少样本识别准确率超越现有最佳方法
  • 适合低资源场景下的跨域推理应用

视觉语言模型在跨模态推理方面取得了显著进展,但现有方法在领域变化时性能下降,或需大量计算资源进行微调。为此,我们提出一种新型大模型适配方法——无需训练的双曲适配器(T-DHA)。该方法将视觉语言概念间的语义关系建模为具有层次树结构的数据,并在双曲空间而非传统欧氏空间中进行表示。双曲空间具有随半径呈指数增长的体积特性,优于欧氏空间的多项式增长。利用庞加莱球模型,该方法在更少特征维度下实现更强的表征与区分能力。结合负样本学习,进一步提升了分类精度与鲁棒性。在多个数据集上的实验表明,T-DHA 在少样本图像识别和领域泛化任务中显著优于现有最先进方法。

原文摘要 · Abstract (English)

Recent research in Vision-Language Models (VLMs) has significantly advanced our capabilities in cross-modal reasoning. However, existing methods suffer from performance degradation with domain changes or require substantial computational resources for fine-tuning in new domains. To address this issue, we develop a new adaptation method for large vision-language models, called \textit{Training-free Dual Hyperbolic Adapters} (T-DHA). We characterize the vision-language relationship between semantic concepts, which typically has a hierarchical tree structure, in the hyperbolic space instead of the traditional Euclidean space. Hyperbolic spaces exhibit exponential volume growth with radius, unlike the polynomial growth in Euclidean space. We find that this unique property is particularly effective for embedding hierarchical data structures using the Poincaré ball model, achieving significantly improved representation and discrimination power. Coupled with negative learning, it provides more accurate and robust classifications with fewer feature dimensions. Our extensive experimental results on various datasets demonstrate that the T-DHA method significantly outperforms existing state-of-the-art methods in few-shot image recognition and domain generalization tasks.

跨模态推理双曲空间零样本适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。