一个模型搞定所有传感器的骨骼数据表示学习,突破跨设备兼容性瓶颈。
One for All: Generalist Foundation Model for Cross-Sensor Skeleton Representation Learning

- 用可学习的通用关节槽+注意力机制统一不同传感器的骨骼结构
- 在10个3D骨骼数据集上实现跨传感器最优性能,超越专用模型
- 适合需要统一处理多源骨骼数据的研究者和开发者
为从大规模无标签数据中学习可泛化的运动表示,自监督学习(SSL)已成为主流方法。然而,现有方法受限于骨骼数据固有的异质性——不同传感器间存在关节数量、索引协议和拓扑结构的差异,通常需为每种传感器或数据集训练独立模型。为此,我们提出SOfA(Skeleton One for All),首个面向跨传感器骨骼表示学习的通用基础模型。为解决因关节数量差异带来的维度鸿沟,引入固定大小的可学习通用关节槽,作为容纳任意骨骼拓扑的统一载体;通过注意力机制动态聚合来自传感器特异性输入的骨骼信息以填充这些槽。此外,通过预训练文本编码器生成语义关节嵌入,而非依赖绝对位置嵌入,解决了不同传感器间的关节索引错位问题。我们对10个3D骨骼数据集进行标准化处理以实现统一训练。大量实验表明,SOfA可作为真正通用的编码器,在多种下游任务和传感器类型上均达到最先进水平,常以单一基础模型超越专门训练的模型。
原文摘要 · Abstract (English)
For learning generalizable motion representations from large-scale unlabeled data, Self-supervised learning (SSL) has become a widely adopted methodology. However, existing approaches are primarily limited by the inherent heterogeneity of skeleton data---characterized by varying joint counts, indexing protocols, and topological structures across different sensors---which typically necessitates training separate, sensor-specific, or even entirely dataset-specific models. To overcome this, we introduce SOfA (Skeleton One for All), the first generalist foundation model designed to achieve sensor-unified skeleton representation learning across diverse sensors. To accommodate the dimensional gap caused by varying joint counts, we introduce a fixed-size set of learnable Canonical Joint Slots, acting as a universal vessel that seamlessly accommodates arbitrary skeletal topologies. SOfA fills these slots via an attention mechanism that dynamically aggregates skeletal information from sensor-specific inputs. Furthermore, we resolve joint index misalignment between various sensors by introducing a Semantic Joint Embedding derived from a pre-trained text encoder, rather than relying on absolute positional embeddings. To validate our approach, we standardized ten 3D skeleton datasets for unified training. Extensive experiments demonstrate that SOfA can serve as a truly universal encoder, achieving state-of-the-art (SOTA) performance across a wide range of downstream tasks and sensor types, often outperforming dataset-specific specialist models with a single foundation model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。