arXiv:2602.02334cs.CVcs.AI2026-02被引 4

用残差向量量化分离动作内容与风格,实现零微调风格迁移。

VQ-Style: Disentangling Style and Content in Motion with Residual Quantized Representations

  • 基于残差向量量化构建粗到细的动作表征层次。
  • 通过对比学习和信息泄露损失增强内容与风格解耦。
  • 推理时直接交换代码块,无需训练即可适配新风格。

人类动作数据蕴含丰富而复杂的语义内容与细微风格特征,难以建模。本文提出一种新型方法,有效解耦动作中的内容与风格,以支持风格迁移。核心思路是:内容对应粗粒度运动属性,风格则捕捉精细表达细节。为此,我们采用残差向量量化变分自编码器(RVQ-VAEs)学习动作的粗到细表征。进一步结合代码本学习、对比学习及新颖的信息泄露损失,将内容与风格在不同代码本中进行组织。利用简单高效的推理阶段技术——量化代码交换,实现无需微调的风格迁移。该框架在多种应用中展现强泛化能力,包括风格迁移、风格去除与动作融合。

原文摘要 · Abstract (English)

Human motion data is inherently rich and complex, containing both semantic content and subtle stylistic features that are challenging to model. We propose a novel method for effective disentanglement of the style and content in human motion data to facilitate style transfer. Our approach is guided by the insight that content corresponds to coarse motion attributes while style captures the finer, expressive details. To model this hierarchy, we employ Residual Vector Quantized Variational Autoencoders (RVQ-VAEs) to learn a coarse-to-fine representation of motion. We further enhance the disentanglement by integrating codebook learning with contrastive learning and a novel information leakage loss to organize the content and the style across different codebooks. We harness this disentangled representation using our simple and effective inference-time technique Quantized Code Swapping, which enables motion style transfer without requiring any fine-tuning for unseen styles. Our framework demonstrates strong versatility across multiple inference applications, including style transfer, style removal, and motion blending.

动作生成风格迁移向量量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。