用多模态输入生成逼真自然的手势,实时性更强。
Large Body Language Models
- 融合文本音频视频输入,用Transformer与扩散模型联合生成动作
- 在FGD上降低30%,在FID上提升25%,表现领先现有方法
- 适合虚拟助手、元宇宙交互等需自然手势的场景
随着虚拟代理在人机交互中日益普及,实现实时、符合语境的逼真手势生成仍面临重大挑战。尽管神经渲染技术在静态脚本上取得进展,但在人机交互中的应用仍受限。为此,我们提出大型肢体语言模型(LBLMs),并设计LBLM-AVA——一种结合Transformer-XL大语言模型与并行扩散模型的新架构,可从多模态输入(文本、音频、视频)生成类人手势。LBLM-AVA包含多项关键组件:多模态到姿态的嵌入、重新定义注意力机制的序列到序列映射、用于动作序列连贯性的时序平滑模块,以及增强真实感的基于注意力的优化模块。模型在我们自有的大规模开源数据集Allo-AVA上训练。相比现有方法,其在生成逼真且语境恰当的手势方面达到顶尖水平,使弗雷谢尔手势距离(FGD)降低30%,弗雷谢尔初始距离(FID)提升25%。
原文摘要 · Abstract (English)
As virtual agents become increasingly prevalent in human-computer interaction, generating realistic and contextually appropriate gestures in real-time remains a significant challenge. While neural rendering techniques have made substantial progress with static scripts, their applicability to human-computer interactions remains limited. To address this, we introduce Large Body Language Models (LBLMs) and present LBLM-AVA, a novel LBLM architecture that combines a Transformer-XL large language model with a parallelized diffusion model to generate human-like gestures from multimodal inputs (text, audio, and video). LBLM-AVA incorporates several key components enhancing its gesture generation capabilities, such as multimodal-to-pose embeddings, enhanced sequence-to-sequence mapping with redefined attention mechanisms, a temporal smoothing module for gesture sequence coherence, and an attention-based refinement module for enhanced realism. The model is trained on our large-scale proprietary open-source dataset Allo-AVA. LBLM-AVA achieves state-of-the-art performance in generating lifelike and contextually appropriate gestures with a 30% reduction in Fréchet Gesture Distance (FGD), and a 25% improvement in Fréchet Inception Distance compared to existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。