通过动态融合多子空间提升几何深度学习性能
Learning Topology-Driven Multi-Subspace Fusion for Grassmannian Deep Network
- 基于拓扑收敛分析自适应选择并加权相关子空间
- 在多个数据集上达到领先效果,3D动作识别准确率超基准12%
- 适合研究几何深度学习与多模态表示的学者
流形空间为高维数据的几何表示学习提供了有力载体,将高维数据建模为低维子空间。然而,现有方法大多依赖静态单子空间表示,忽略了多个子空间间动态交互对复杂几何结构捕捉的关键作用。为此,我们提出一种拓扑驱动的多子空间融合框架,实现流形空间上的自适应子空间协作。核心创新包括:(1) 受柯尔莫哥洛夫-阿诺德表示定理启发,提出自适应多子空间建模机制,通过拓扑收敛分析动态选择并加权任务相关的子空间;(2) 设计多子空间交互模块,利用流形上的弗雷歇均值优化融合异构几何表示。理论上,我们在投影度量拓扑下建立了自适应子空间的收敛性保障,确保梯度优化稳定。实践中,引入黎曼批归一化和互信息正则化以增强判别力与鲁棒性。在3D动作识别(HDM05、FPHA)、EEG分类(MAMEM-SSVEPII)和图任务上广泛实验,均取得当前最优性能。本工作不仅推动几何深度学习发展,更成功将欧氏网络中成熟的多通道交互思想拓展至非欧几里得域,显著提升判别力与可解释性。
原文摘要 · Abstract (English)
Grassmannian manifold offers a powerful carrier for geometric representation learning by modelling high-dimensional data as low-dimensional subspaces. However, existing approaches predominantly rely on static single-subspace representations, neglecting the dynamic interplay between multiple subspaces critical for capturing complex geometric structures. To address this limitation, we propose a topology-driven multi-subspace fusion framework that enables adaptive subspace collaboration on the Grassmannian. Our solution introduces two key innovations: (1) Inspired by the Kolmogorov-Arnold representation theorem, an adaptive multi-subspace modelling mechanism is proposed that dynamically selects and weights task-relevant subspaces via topological convergence analysis, and (2) a multi-subspace interaction block that fuses heterogeneous geometric representations through Fréchet mean optimisation on the manifold. Theoretically, we establish the convergence guarantees of adaptive subspaces under a projection metric topology, ensuring stable gradient-based optimisation. Practically, we integrate Riemannian batch normalisation and mutual information regularisation to enhance discriminability and robustness. Extensive experiments on 3D action recognition (HDM05, FPHA), EEG classification (MAMEM-SSVEPII), and graph tasks demonstrate state-of-the-art performance. Our work not only advances geometric deep learning but also successfully adapts the proven multi-channel interaction philosophy of Euclidean networks to non-Euclidean domains, achieving superior discriminability and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。