用3D控制信号实现高精度手物交互视频生成,打通真实与仿真数据壁垒。
HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

- 基于3D ControlNet的扩散模型,显式编码几何与运动信息。
- 在TASTE-Rob上实现最优空间保真度与时间连贯性。
- 支持真实/仿真3D信号输入,适合需要精准控制的场景。
近期方法在手物交互视频的视觉质量上取得显著进展,但多数依赖缺乏空间表达力的2D控制信号,限制了对合成3D条件数据的利用。为此,我们提出HVG-3D,一个面向显式3D表示的3D感知手物交互(HOI)视频生成统一框架。具体地,设计了一种结合3D ControlNet的扩散架构,通过编码3D输入中的几何与运动线索,实现生成过程中的显式3D推理。HVG-3D包含两个核心组件:(i) 3D感知的HOI视频生成扩散架构,用于从3D输入中编码几何与运动线索以实现显式3D推理;(ii) 混合管道构建输入与条件信号,支持训练与推理阶段的灵活精确控制。推理时,给定单张真实图像和来自仿真或真实数据的3D控制信号,HVG-3D可生成高保真、时间一致的视频,并具备精确的空间与时间控制能力。在TASTE-Rob数据集上的实验表明,HVG-3D在空间保真度、时间连贯性和可控性方面均达到当前最优水平,同时有效利用真实与仿真数据。
原文摘要 · Abstract (English)
Recent methods have made notable progress in the visual quality of hand-object interaction video synthesis. However, most approaches rely on 2D control signals that lack spatial expressiveness and limit the utilization of synthetic 3D conditional data. To address these limitations, we propose HVG-3D, a unified framework for 3D-aware hand-object interaction (HOI) video synthesis conditioned on explicit 3D representations. Specifically, we develop a diffusion-based architecture augmented with a 3D ControlNet, which encodes geometric and motion cues from 3D inputs to enable explicit 3D reasoning during video synthesis. To achieve high-quality synthesis, HVG-3D is designed with two core components: (i) a 3D-aware HOI video generation diffusion architecture that encodes geometric and motion cues from 3D inputs for explicit 3D reasoning; and (ii) a hybrid pipeline for constructing input and condition signals, enabling flexible and precise control during both training and inference. During inference, given a single real image and a 3D control signal from either simulation or real data, HVG-3D generates high-fidelity, temporally consistent videos with precise spatial and temporal control. Experiments on the TASTE-Rob dataset demonstrate that HVG-3D achieves state-of-the-art spatial fidelity, temporal coherence, and controllability, while enabling effective utilization of both real and simulated data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。