构建大规模多模态对话数据集,实现虚拟角色动作与语音精准同步
Allo-AVA: A Large-Scale Multimodal Conversational AI Dataset for Allocentric Avatar Gesture Animation
- 基于第三人称视角,整合语音、文本与动作关键点
- 包含约1250小时视频,关键点精确对齐时间戳
- 适合虚拟现实、数字助手等场景的自然动作生成研究
高质量多模态训练数据的匮乏严重制约了虚拟环境中对话式人工智能角色的拟人化动画生成。现有数据集常缺乏语音、面部表情与肢体动作之间的精细同步,难以体现真实人类交流特征。为此,我们提出 Allo-AVA,一个面向第三人称视角(allocentric)下文本和音频驱动角色动作动画的大规模数据集。Allo-AVA 包含约 1,250 小时多样化的视频内容,配有音频、转录文本及提取的关键点数据。其独特之处在于将关键点精确映射到时间戳,实现语音与身体及面部动作的高度同步再现。该资源可有效推动更自然、上下文感知的虚拟角色动画模型的研发与评估,有望在虚拟现实、数字助理等领域带来变革。
原文摘要 · Abstract (English)
The scarcity of high-quality, multimodal training data severely hinders the creation of lifelike avatar animations for conversational AI in virtual environments. Existing datasets often lack the intricate synchronization between speech, facial expressions, and body movements that characterize natural human communication. To address this critical gap, we introduce Allo-AVA, a large-scale dataset specifically designed for text and audio-driven avatar gesture animation in an allocentric (third person point-of-view) context. Allo-AVA consists of $\sim$1,250 hours of diverse video content, complete with audio, transcripts, and extracted keypoints. Allo-AVA uniquely maps these keypoints to precise timestamps, enabling accurate replication of human movements (body and facial gestures) in synchronization with speech. This comprehensive resource enables the development and evaluation of more natural, context-aware avatar animation models, potentially transforming applications ranging from virtual reality to digital assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。