让虚拟人同时听懂话、用上物体,生成自然动作。
InteracTalker: Prompt-Based Human-Object Interaction with Co-Speech Gesture Generation
- 用提示词驱动,统一建模说话与物体交互动作
- 在真实数据集上生成更逼真、可控的全身动作
- 适合做虚拟助手、数字人动画的开发者
生成能自然响应语言和物理对象的人体动作对交互式数字体验至关重要。现有方法分别处理语音驱动手势或物体交互,因缺乏整合性数据集而限制实际应用。为此,我们提出 InteracTalker,一个将提示词驱动的物体感知交互与伴随言语动作生成无缝融合的新框架。通过多阶段训练学习统一的动作、语音与提示嵌入空间,并构建了一个丰富的人-物交互数据集,由现有文本到动作数据集扩展而来,包含详细的物体交互标注。框架采用通用动作适配模块,可独立训练并动态组合于推理阶段。为解决异构条件信号间的不平衡问题,提出自适应融合策略,在扩散采样中动态重加权条件信号。InteracTalker 成功统一此前分离的任务,在言语伴随动作生成与物体交互合成方面均优于先前方法,生成高度真实、具备物体感知的全身动作,显著提升真实感、灵活性与控制力。
原文摘要 · Abstract (English)
Generating realistic human motions that naturally respond to both spoken language and physical objects is crucial for interactive digital experiences. Current methods, however, address speech-driven gestures or object interactions independently, limiting real-world applicability due to a lack of integrated, comprehensive datasets. To overcome this, we introduce InteracTalker, a novel framework that seamlessly integrates prompt-based object-aware interactions with co-speech gesture generation. We achieve this by employing a multi-stage training process to learn a unified motion, speech, and prompt embedding space. To support this, we curate a rich human-object interaction dataset, formed by augmenting an existing text-to-motion dataset with detailed object interaction annotations. Our framework utilizes a Generalized Motion Adaptation Module that enables independent training, adapting to the corresponding motion condition, which is then dynamically combined during inference. To address the imbalance between heterogeneous conditioning signals, we propose an adaptive fusion strategy, which dynamically reweights the conditioning signals during diffusion sampling. InteracTalker successfully unifies these previously separate tasks, outperforming prior methods in both co-speech gesture generation and object-interaction synthesis, outperforming gesture-focused diffusion methods, yielding highly realistic, object-aware full-body motions with enhanced realism, flexibility, and control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。