arXiv:2603.04338cs.CV2026-03

用视频生成3D人体与可动物体交互,不依赖真实3D数据。

ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors

  • 从单目视频出发,通过逆向渲染重建4D动态场景。
  • 在打开冰箱等动作中接触准确率显著提升,穿透减少60%以上。
  • 适合做物理合理的人体-物体交互生成,如动画、游戏开发。

无3D/4D标注条件下,合成物理合理的可动人体-物体交互(HOI)仍是根本挑战。现有零样本方法虽利用视频扩散模型生成交互,但主要局限于刚性物体操作,缺乏显式的4D几何推理。为此,本文将可动HOI生成建模为从单目视频先验出发的4D重建问题:仅需扩散模型生成的视频,即可重建完整4D可动场景,无需任何3D监督。该重建方法将生成的2D视频作为逆向渲染的监督信号,恢复几何一致且物理合理的4D场景,自然满足接触、关节运动和时序一致性。提出ArtHOI,首个基于视频先验的4D重建式零样本可动人体-物体交互生成框架。核心设计包括:1)基于光流的部件分割:利用光流区分视频中动态与静态区域;2)解耦重建流程:由于单目模糊性导致人体运动与物体关节联合优化不稳定,先恢复物体关节状态,再条件化生成人体运动。ArtHOI融合视频生成与几何感知重建,生成语义对齐且物理可信的交互。在多种可动场景(如开冰箱、柜子、微波炉)中,相比先前方法,接触准确率更高,穿透减少60%以上,关节还原度显著提升,首次实现超越刚性操作的零样本可动交互生成。

原文摘要 · Abstract (English)

Synthesizing physically plausible articulated human-object interactions (HOI) without 3D/4D supervision remains a fundamental challenge. While recent zero-shot approaches leverage video diffusion models to synthesize human-object interactions, they are largely confined to rigid-object manipulation and lack explicit 4D geometric reasoning. To bridge this gap, we formulate articulated HOI synthesis as a 4D reconstruction problem from monocular video priors: given only a video generated by a diffusion model, we reconstruct a full 4D articulated scene without any 3D supervision. This reconstruction-based approach treats the generated 2D video as supervision for an inverse rendering problem, recovering geometrically consistent and physically plausible 4D scenes that naturally respect contact, articulation, and temporal coherence. We introduce ArtHOI, the first zero-shot framework for articulated human-object interaction synthesis via 4D reconstruction from video priors. Our key designs are: 1) Flow-based part segmentation: leveraging optical flow as a geometric cue to disentangle dynamic from static regions in monocular video; 2) Decoupled reconstruction pipeline: joint optimization of human motion and object articulation is unstable under monocular ambiguity, so we first recover object articulation, then synthesize human motion conditioned on the reconstructed object states. ArtHOI bridges video-based generation and geometry-aware reconstruction, producing interactions that are both semantically aligned and physically grounded. Across diverse articulated scenes (e.g., opening fridges, cabinets, microwaves), ArtHOI significantly outperforms prior methods in contact accuracy, penetration reduction, and articulation fidelity, extending zero-shot interaction synthesis beyond rigid manipulation through reconstruction-informed synthesis.

人机交互4D重建视频生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。