统一生成语音与手势,让表达更自然同步。
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
- 用交错的离散标记序列联合建模语音和手势
- 手势生成质量优于单一模态基线模型
- 支持多说话人、多风格克隆及仅凭语音生成手势
人类交流是多模态的,语音与手势紧密耦合,但现有大多数语音与手势生成方法采用顺序合成,削弱了同步性和语调对齐。我们提出Gelina,一个统一框架,通过在离散自回归主干中使用交错标记序列,联合从文本生成语音与伴随手势,并采用模态专用解码器。Gelina支持多说话人和多风格克隆,且可实现仅基于语音输入生成手势。主观与客观评估表明,其语音质量具有竞争力,手势生成效果优于单模态基线。
原文摘要 · Abstract (English)
Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-speech gestures from text using interleaved token sequences in a discrete autoregressive backbone, with modality-specific decoders. Gelina supports multi-speaker and multi-style cloning and enables gesture-only synthesis from speech inputs. Subjective and objective evaluations demonstrate competitive speech quality and improved gesture generation over unimodal baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。