arXiv:2603.20187cs.CV2026-03

让视频生成更自然的人类反应动作,解决视觉与行为不匹配问题

MuSteerNet: Human Reaction Generation from Videos via Observation-Reaction Mutual Steering

  • 通过观察与反应相互引导机制,修正视觉输入与动作之间的关系偏差
  • 在Human3.6M和MPII-React数据集上达到领先性能,动作匹配度显著提升
  • 适合需要逼真人机交互的虚拟助手、动画生成等场景

视频驱动的人类反应生成旨在合成直接响应视频序列的3D人体动作,对构建类人交互式AI系统至关重要。然而现有方法难以有效利用视频输入来引导反应生成,导致动作与视频内容不匹配。我们发现这一局限源于视觉观察与反应类型间的严重关系扭曲。为此,提出MuSteerNet框架,通过观察-反应互驱机制生成3D人类反应。首先设计原型反馈引导机制,利用门控增量修正模块与关系边界约束,基于从人类反应中学习的原型向量优化视觉观察。随后引入双耦合反应精炼模块,充分使用修正后的视觉线索进一步优化生成动作,显著提升反应质量。大量实验与消融研究验证了方法的有效性。代码即将发布:https://github.com/zhouyuan888888/MuSteerNet。

原文摘要 · Abstract (English)

Video-driven human reaction generation aims to synthesize 3D human motions that directly react to observed video sequences, which is crucial for building human-like interactive AI systems. However, existing methods often fail to effectively leverage video inputs to steer human reaction synthesis, resulting in reaction motions that are mismatched with the content of video sequences. We reveal that this limitation arises from a severe relational distortion between visual observations and reaction types. In light of this, we propose MuSteerNet, a simple yet effective framework that generates 3D human reactions from videos via observation-reaction mutual steering. Specifically, we first propose a Prototype Feedback Steering mechanism to mitigate relational distortion by refining visual observations with a gated delta-rectification modulator and a relational margin constraint, guided by prototypical vectors learned from human reactions. We then introduce Dual-Coupled Reaction Refinement that fully leverages rectified visual cues to further steer the refinement of generated reaction motions, thereby effectively improving reaction quality and enabling MuSteerNet to achieve competitive performance. Extensive experiments and ablation studies validate the effectiveness of our method. Code coming soon: https://github.com/zhouyuan888888/MuSteerNet.

动作生成视频理解人机交互3D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。