让虚拟人根据文字指令与物体互动,生成更自然的说话虚拟形象。
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
- 分两路处理:感知环境+规划动作,再合成视频,提升可控性。
- 在新基准上生成动作与物体交互更对齐文本描述,视觉质量更高。
- 适合做虚拟主播、数字人交互系统的研究者或开发者参考。
生成说话虚拟人是视频生成的基础任务。尽管现有方法能生成带有简单人体动作的全身说话虚拟人,但将其扩展到基于环境的真人-物体交互(GHOI)仍是开放挑战,要求虚拟人根据文本指令与周围物体进行对齐交互。该挑战源于环境感知需求和生成中控制力与质量之间的权衡问题。为此,我们提出一种新型双流框架InteractAvatar,将环境感知与动作规划从视频合成中解耦。通过目标检测增强环境感知,引入感知与交互模块(PIM)生成对齐文本的交互动作;同时设计音频-交互感知生成模块(AIM),合成生动的交互式说话虚拟人视频。通过专门设计的动作-视频对齐器,PIM与AIM共享相似网络结构,实现动作与合理视频的并行联合生成,有效缓解控制-质量困境。最后,我们构建了用于评估GHOI视频生成的新基准GroundedInter。大量实验与对比验证了本方法在生成接地型真人-物体交互说话虚拟人方面的有效性。
原文摘要 · Abstract (English)
Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open challenge, requiring the avatar to perform text-aligned interactions with surrounding objects. This challenge stems from the need for environmental perception and the control-quality dilemma in GHOI generation. To address this, we propose a novel dual-stream framework, InteractAvatar, which decouples perception and planning from video synthesis for grounded human-object interaction. Leveraging detection to enhance environmental perception, we introduce a Perception and Interaction Module (PIM) to generate text-aligned interaction motions. Additionally, an Audio-Interaction Aware Generation Module (AIM) is proposed to synthesize vivid talking avatars performing object interactions. With a specially designed motion-to-video aligner, PIM and AIM share a similar network structure and enable parallel co-generation of motions and plausible videos, effectively mitigating the control-quality dilemma. Finally, we establish a benchmark, GroundedInter, for evaluating GHOI video generation. Extensive experiments and comparisons demonstrate the effectiveness of our method in generating grounded human-object interactions for talking avatars. Project page: https://interactavatar.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。