arXiv:2605.07252cs.GRcs.CV2026-05被引 1

让语音生成个性化手势,既懂语义又保持风格一致。

PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation

论文配图:PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
图 1 · 摘自论文原文
  • 用语义引导的分层编码,分离动作内容与个人风格
  • 在多个数据集上实现领先性能,风格一致性显著提升
  • 适合需要定制化虚拟角色动作的应用场景

对话语音手势生成旨在合成与语音语义一致且符合用户指定风格的真实身体动作。现有基于VQ-VAE的方法虽提升了生成质量,但未能将语义结构融入动作表征,也未显式解耦内容与风格,限制了语义连贯性与个性化精度。本文提出PersonaGest,一种两阶段框架:第一阶段采用语义引导的残差向量量化自编码器(RVQ-VAE),在残差量化结构中解耦动作内容与手势风格,其中语义感知动作码本(SMoC)按手势语义组织内容码本,对比学习进一步强化内容-风格分离;第二阶段通过掩码生成式变压器以语义感知重掩码策略生成内容标记,并由一系列风格残差变换器根据参考动作提示进行风格控制。大量实验表明,在客观指标与主观用户评估中均达到当前最优表现,且对参考提示的风格一致性强。项目主页含演示视频:https://danny-nus.github.io/PersonaGest/

原文摘要 · Abstract (English)

Co-speech gesture generation aims to synthesize realistic body movements that are semantically coherent with speech and faithful to a user-specified gestural style. Existing VQ-VAE based co-speech gesture generation methods improve generation quality but fail to encode semantic structure into the motion representation or explicitly disentangle content from style, limiting both semantic coherence and personalization fidelity. We present PersonaGest, a two-stage framework addressing both limitations. In the first stage, a semantic-guided RVQ-VAE disentangles motion content and gestural style within the residual quantization structure, where a Semantic-Aware Motion Codebook (SMoC) organizes the content codebook by gesture semantics and contrastive learning further enforces content-style separation. In the second stage, a Masked Generative Transformer generates content tokens via a semantic-aware re-masking strategy, followed by a cascade of Style Residual Transformers conditioned on a reference motion prompt for style control. Extensive experiments demonstrate state-of-the-art performance on objective metrics and perceptual user studies, with strong style consistency to the reference prompt. Our project page with demo videos is available at https://danny-nus.github.io/PersonaGest/

手势生成个性化语义控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。