用双曲几何提升图文数据蒸馏,让模型更省力、更鲁棒。
Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

- 将图文表示映射到双曲空间,分层对齐共享语义
- 在固定预算下,检索与迁移性能优于现有方法
- 适合资源受限场景下的多模态模型训练
视觉-语言数据蒸馏(VLDD)将大规模图像-文本配对数据压缩为少量合成样本,以在严苛的数据与计算预算下高效训练对比式多模态模型。现有方法多通过匹配专家轨迹或跨模态统计信息进行对齐,但仍在欧氏空间中强制全维对齐,易因图像-文本相关性秩不足而失效——共享语义集中于低维子空间,剩余变化分散在弱相关残差空间。LoRS通过低秩分解放松相似度对齐,但未显式控制表示空间中的主导对齐能力与结构。为此,我们提出一种秩感知双曲对齐(RAHA),结合层次几何与显式对齐容量控制:将多模态表示提升至双曲空间,采用非对称目标函数,在共享范围强制测地线对齐,同时正则化残差空间以保留模态私有多样性,增强迁移鲁棒性。在多个基准测试中,RAHA在固定预算下展现出竞争力的跨模态检索表现及更优的迁移指标。
原文摘要 · Abstract (English)
Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can efficiently train contrastive vision-language models under strict data and compute budgets. Most existing methods match expert trajectories or cross-modal statistics, yet still enforce full-dimensional alignment in a Euclidean embedding space. This is often overly restrictive due to rank-deficient image--text correlation, with shared semantics concentrated in a low-dimensional range and remaining variation spread across a weakly correlated residual subspace. LoRS relaxes alignment at the similarity level by low-rank factorization, but does not explicitly control dominant alignment capacity and structure in the representation space. We thus propose a rank-aware hyperbolic alignment (RAHA) that combines hierarchical geometry with explicit alignment-capacity control. RAHA lifts multimodal representations to hyperbolic space and optimizes distilled pairs with asymmetric objectives that enforce geodesic alignment in the shared range while regularizing the residual subspace to preserve modality-private diversity and improve transfer robustness. Experiments on benchmarks show that RAHA demonstrates competitive cross-modal retrieval and improved transfer indicators under fixed budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。