arXiv:2503.08270cs.CV2025-03ICCV被引 15

从视频生成3D人类反应,涵盖人-人、人-动物、人-场景互动。

HERO: Human Reaction Generation from Videos

  • 融合全局与帧级局部视觉特征,捕捉交互意图
  • 在ViMo数据集上生成反应更自然,覆盖三类互动
  • 适合做交互式AI、虚拟人、动作合成的研究者

人类反应生成是互动AI的重要研究方向,因人类持续与环境互动。以往工作主要基于人体运动序列生成反应,限制于人-人交互,且忽略情绪影响。本文提出HERO框架,从RGB视频生成3D人类反应,涵盖人-人、人-动物及人-场景互动,天然包含表情等情绪信息。HERO通过融合全局与帧级局部视觉表示提取交互意图,并以此指导反应生成;同时持续注入局部视觉特征,充分利用视频动态特性。为支持该任务,构建了包含配对视频-动作数据的ViMo数据集。大量实验表明方法优越性。代码与数据集将公开于https://jackyu6.github.io/HERO。

原文摘要 · Abstract (English)

Human reaction generation represents a significant research domain for interactive AI, as humans constantly interact with their surroundings. Previous works focus mainly on synthesizing the reactive motion given a human motion sequence. This paradigm limits interaction categories to human-human interactions and ignores emotions that may influence reaction generation. In this work, we propose to generate 3D human reactions from RGB videos, which involves a wider range of interaction categories and naturally provides information about expressions that may reflect the subject's emotions. To cope with this task, we present HERO, a simple yet powerful framework for Human rEaction geneRation from videOs. HERO considers both global and frame-level local representations of the video to extract the interaction intention, and then uses the extracted interaction intention to guide the synthesis of the reaction. Besides, local visual representations are continuously injected into the model to maximize the exploitation of the dynamic properties inherent in videos. Furthermore, the ViMo dataset containing paired Video-Motion data is collected to support the task. In addition to human-human interactions, these video-motion pairs also cover animal-human interactions and scene-human interactions. Extensive experiments demonstrate the superiority of our methodology. The code and dataset will be publicly available at https://jackyu6.github.io/HERO.

视频生成动作合成交互理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。