arXiv:2603.24383cs.CV2026-03中稿 · CVPR被引 1

用2D图像提取视觉先验,提升3D人物交互生成真实感。

ViHOI: Human-Object Interaction Synthesis with Visual Priors

  • 从2D图像中提取视觉与文本先验,指导3D动作生成。
  • 在多个基准上超越现有方法,对未见物体/动作泛化能力强。
  • 适合需要高真实感交互生成的研究与应用者。

生成真实且符合物理规律的3D人物-物体交互(HOI)仍是动作生成中的关键挑战。主要原因在于仅靠文字难以准确描述这些物理约束。为此,我们提出新范式:从易获取的2D图像中提取丰富的交互先验。具体而言,提出ViHOI框架,使基于扩散模型的生成器能利用来自2D图像的任务特定先验,提升生成质量。我们采用大型视觉语言模型(VLM)作为强大的先验提取引擎,并采用层解耦策略获取视觉与文本先验。同时,设计基于Q-Former的适配器,将VLM的高维特征压缩为紧凑的先验标记,显著促进扩散模型的条件训练。框架在运动渲染图像数据集上训练,确保视觉输入与动作序列间的严格语义对齐。推理时,利用文生图模型合成参考图像,增强对未见物体和交互类别的泛化能力。实验表明,ViHOI在多个基准上达到最先进性能,展现出优越泛化能力。

原文摘要 · Abstract (English)

Generating realistic and physically plausible 3D Human-Object Interactions (HOI) remains a key challenge in motion generation. One primary reason is that describing these physical constraints with words alone is difficult. To address this limitation, we propose a new paradigm: extracting rich interaction priors from easily accessible 2D images. Specifically, we introduce ViHOI, a novel framework that enables diffusion-based generative models to leverage rich, task-specific priors from 2D images to enhance generation quality. We utilize a large Vision-Language Model (VLM) as a powerful prior-extraction engine and adopt a layer-decoupled strategy to obtain visual and textual priors. Concurrently, we design a Q-Former-based adapter that compresses the VLM's high-dimensional features into compact prior tokens, which significantly facilitates the conditional training of our diffusion model. Our framework is trained on motion-rendered images from the dataset to ensure strict semantic alignment between visual inputs and motion sequences. During inference, it leverages reference images synthesized by a text-to-image generation model to improve generalization to unseen objects and interaction categories. Experimental results demonstrate that ViHOI achieves state-of-the-art performance, outperforming existing methods across multiple benchmarks and demonstrating superior generalization.

3D生成交互合成视觉先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。