arXiv:2608.10857cs.LG2026-08

FiGuRO可自动估算多模态数据内在维度,实现高效解耦表征。

FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data

论文配图:FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data
图 1 · 摘自论文原文
  • 基于截断SVD与动态秩优化,自适应确定低秩投影维数
  • 在模拟与真实数据中准确捕捉不同子空间尺度和共享/私有信息比例
  • 无需复杂损失函数,适用于预训练模型的后处理解耦

确定数据复杂度(即内在维度,ID)是实现高效可解释表征学习的基础。在多模态场景下,学习共享与私有信息的解耦表示尤为困难。现有方法存在关键缺陷:多为静态、单模态,或对比方法仅隐式适应共享维度。我们提出保真度引导的秩优化框架(FiGuRO),在模型容量与超参数约束下,近似单模态与多模态数据的内在维度。FiGuRO通过截断奇异值分解学习低秩投影维数,并设计算法决定何时增减维度及在哪个潜在空间操作。共享与私有信息的解耦作为优化的涌现属性自然产生,无需复杂辅助损失函数。实验表明,FiGuRO优于现有ID估计方法,对超参数变化更鲁棒。在模拟与真实数据中,能准确捕捉不同内在维度尺度和子空间比例,并成功分解共享与私有信息。此外,该方法可应用于现代单模态预训练模型,实现多模态表征的高效后处理解耦。

原文摘要 · Abstract (English)

Determining the complexity, or Intrinsic Dimension (ID), of data is fundamental to efficient and interpretable representation learning. This is particularly challenging in multi-modal settings when trying to learn disentangled representations for shared and private information. Existing techniques leave a critical gap: they are often static, uni-modal, or in the case of contrastive methods, adapt only to the shared ID implicitly. We introduce Fidelity-Guided Rank Optimization (FiGuRO), a framework for approximating the ID of uni- and multi-modal data under constraints of model capacity and hyperparameters. FiGuRO learns the dimensions of low-rank projections using truncated singular value decomposition and an algorithm that determines when to reduce or increase dimension and in which latent space. Disentanglement of shared and private information arises as an emergent property of this optimization, eliminating the need for complex auxiliary loss functions. We demonstrate that FiGuRO outperforms existing ID estimation techniques and is more robust to hyperparameter changes. Across simulations and real-world data, FiGuRO captures distinct ID scales and varying subspace ratios, and decomposes shared and private information successfully. Furthermore, we show that FiGuRO can be applied to modern uni-modal pretrained models, enabling efficient, post-hoc disentanglement of multi-modal representations.

内在维度多模态解耦表征降维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。