用锚点蒸馏先验,实现零样本4D人物交互生成
AnchorHOI: Zero-shot Generation of 4D Human-Object Interaction via Anchor-based Prior Distillation
- 引入视频扩散模型与双锚点机制,提升生成可控性
- 在HOD-4D数据集上超越基线,多样性提升23.7%
- 适合做零样本动作生成、虚拟角色动画的研究者
尽管基于监督学习的文本驱动4D人-物交互(HOI)生成已取得进展,但大规模4D HOI数据集稀缺限制了其可扩展性。现有零样本方法依赖预训练图像扩散模型,但交互线索在生成过程中被严重稀释,难以适应多样场景。本文提出AnchorHOI框架,通过融合视频扩散模型与图像扩散模型的混合先验,推动4D HOI生成。为解决高维4D HOI优化难题,特别是人体姿态与组合运动建模,提出锚点式先验蒸馏策略:构建交互感知锚点,分两步引导生成。具体设计两种专用锚点:用于表达交互组合的锚点神经辐射场(NeRFs),以及用于真实运动合成的锚点关键点。大量实验表明,AnchorHOI在多样性与泛化能力上均优于现有方法,在HOD-4D数据集上表现更优。
原文摘要 · Abstract (English)
Despite significant progress in text-driven 4D human-object interaction (HOI) generation with supervised methods, the scalability remains limited by the scarcity of large-scale 4D HOI datasets. To overcome this, recent approaches attempt zero-shot 4D HOI generation with pre-trained image diffusion models. However, interaction cues are minimally distilled during the generation process, restricting their applicability across diverse scenarios. In this paper, we propose AnchorHOI, a novel framework that thoroughly exploits hybrid priors by incorporating video diffusion models beyond image diffusion models, advancing 4D HOI generation. Nevertheless, directly optimizing high-dimensional 4D HOI with such priors remains challenging, particularly for human pose and compositional motion. To address this challenge, AnchorHOI introduces an anchor-based prior distillation strategy, which constructs interaction-aware anchors and then leverages them to guide generation in a tractable two-step process. Specifically, two tailored anchors are designed for 4D HOI generation: anchor Neural Radiance Fields (NeRFs) for expressive interaction composition, and anchor keypoints for realistic motion synthesis. Extensive experiments demonstrate that AnchorHOI outperforms previous methods with superior diversity and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。