用上下文学习统一建模跨域3D人体动作,提升泛化能力。
Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning
- 基于上下文学习构建统一模型,避免分阶段训练和领域专用模块。
- 在多个数据集和任务上超越前代模型,支持姿态与网格表示融合。
- 适合需要跨模态、跨任务部署的3D人体动作研究者使用。
本文旨在实现跨域3D人体动作建模,要求单一模型能处理多种模态、任务和数据集。现有跨域模型常依赖领域特定组件和多阶段训练,限制了实用性与可扩展性。为此,我们提出一种新训练范式,通过单一流程训练统一跨域模型,消除领域专用组件与多阶段训练需求。首先引入姿势-上下文(Pose-in-Context, PiC),利用上下文学习构建以姿态为中心的跨域模型。尽管PiC在多种基于姿态的任务与数据集上具备泛化能力,但在处理模态多样性、策略变化及上下文依赖时仍存挑战。因此,我们提出人类-上下文(Human-in-Context, HiC),作为PiC的扩展,进一步提升在模态、任务与数据集上的泛化能力。HiC在统一框架中融合姿态与网格表示,扩大任务覆盖范围,并引入更大规模数据集。此外,HiC采用最大最小相似性提示采样策略以增强跨域泛化,设计双分支上下文注入网络架构以更好处理上下文依赖。大量实验表明,相较于PiC,HiC在泛化能力、数据规模适应性和多领域性能方面均表现更优。结果证明,HiC为构建更具灵活性与可扩展性的统一跨域3D人体动作模型提供了可行路径。源代码与模型已公开于https://github.com/BradleyWang0416/Human-in-Context。
原文摘要 · Abstract (English)
This paper aims to model 3D human motion across domains, where a single model is expected to handle multiple modalities, tasks, and datasets. Existing cross-domain models often rely on domain-specific components and multi-stage training, which limits their practicality and scalability. To overcome these challenges, we propose a new setting to train a unified cross-domain model through a single process, eliminating the need for domain-specific components and multi-stage training. We first introduce Pose-in-Context (PiC), which leverages in-context learning to create a pose-centric cross-domain model. While PiC generalizes across multiple pose-based tasks and datasets, it encounters difficulties with modality diversity, prompting strategy, and contextual dependency handling. We thus propose Human-in-Context (HiC), an extension of PiC that broadens generalization across modalities, tasks, and datasets. HiC combines pose and mesh representations within a unified framework, expands task coverage, and incorporates larger-scale datasets. Additionally, HiC introduces a max-min similarity prompt sampling strategy to enhance generalization across diverse domains and a network architecture with dual-branch context injection for improved handling of contextual dependencies. Extensive experimental results show that HiC performs better than PiC in terms of generalization, data scale, and performance across a wide range of domains. These results demonstrate the potential of HiC for building a unified cross-domain 3D human motion model with improved flexibility and scalability. The source codes and models are available at https://github.com/BradleyWang0416/Human-in-Context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。