arXiv:2603.14686cs.CVcs.AI2026-03被引 2

用3D基础模型实现多视角下复杂人物交互视频重演

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model

  • 通过隐式运动描述与多视角推理生成物体姿态序列
  • 在复杂视角变化下保持动作一致性,显著提升交互真实感
  • 适合需要高保真人物-物体交互的视频生成研究者

人-物交互(HOI)视频重演旨在将源视频中的交互动态迁移到新物体上,同时保持手物协调的真实感。现有方法依赖稀疏2D运动控制和单目参考,难以应对复杂非平面运动和大视角变化。本文提出MVHOI,一个两阶段框架:第一阶段,运动提取器将物体动态编码为隐式运动描述符;基于这些描述符,运动驱动物体先验(MDOP)模块调用3D基础模型,通过多视角参考自回归预测粗略物体锚点,即追踪物体在源运动下的姿态与外观演变序列,无需显式姿态估计。第二阶段,基于DiT的视频生成模型以这些锚点为结构引导,多视角参考为外观引导,并复用MDOP的跨视角注意力作为软注意力偏置,减少参考视图混淆。针对长视频,采用交叉迭代推理策略,利用优化后的视频输出刷新后续物体先验。实验表明,该方法在物体保真度、动作一致性、视觉质量与交互真实感方面均优于现有最优方法。

原文摘要 · Abstract (English)

Human-Object Interaction (HOI) video reenactment aims to transfer the interaction dynamics of a source video to a novel target object while preserving realistic hand-object coordination. Existing methods typically rely on sparse 2D motion controls and monocular references, which are insufficient for complex out-of-plane motion and large viewpoint changes. We present MVHOI, a two-stage framework combining implicit motion extraction, 3D-aware multi-view reasoning, and video generation. In the first stage, a motion extractor encodes object dynamics into implicit motion descriptors. Conditioned on these descriptors, our Motion-Driven Object Prior (MDOP) module queries a 3D foundation model over multi-view references of the target object and autoregressively predicts coarse object anchors, a sequence of images that track the object's evolving orientation and appearance under the source motion without any explicit pose estimation. In the second stage, a DiT-based video generation model uses these anchors as structural guidance and the multi-view references as appearance guidance. We further reuse cross-view attention from MDOP as a soft attention bias to reduce reference-view confusion. For long videos, a cross-iterative inference strategy refreshes subsequent object priors using refined video outputs. Experiments demonstrate consistent improvements over state-of-the-art methods in object fidelity, motion consistency, visual quality, and interaction realism.

视频重演3D生成多视角人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。