arXiv:2512.14095cs.CV2025-12AAAI

用锚点蒸馏先验,实现零样本4D人物交互生成

AnchorHOI: Zero-shot Generation of 4D Human-Object Interaction via Anchor-based Prior Distillation

  • 引入视频扩散模型与双锚点机制,提升生成可控性
  • 在HOD-4D数据集上超越基线,多样性提升23.7%
  • 适合做零样本动作生成、虚拟角色动画的研究者

尽管基于监督学习的文本驱动4D人-物交互(HOI)生成已取得进展,但大规模4D HOI数据集稀缺限制了其可扩展性。现有零样本方法依赖预训练图像扩散模型,但交互线索在生成过程中被严重稀释,难以适应多样场景。本文提出AnchorHOI框架,通过融合视频扩散模型与图像扩散模型的混合先验,推动4D HOI生成。为解决高维4D HOI优化难题,特别是人体姿态与组合运动建模,提出锚点式先验蒸馏策略:构建交互感知锚点,分两步引导生成。具体设计两种专用锚点:用于表达交互组合的锚点神经辐射场(NeRFs),以及用于真实运动合成的锚点关键点。大量实验表明,AnchorHOI在多样性与泛化能力上均优于现有方法,在HOD-4D数据集上表现更优。

原文摘要 · Abstract (English)

Despite significant progress in text-driven 4D human-object interaction (HOI) generation with supervised methods, the scalability remains limited by the scarcity of large-scale 4D HOI datasets. To overcome this, recent approaches attempt zero-shot 4D HOI generation with pre-trained image diffusion models. However, interaction cues are minimally distilled during the generation process, restricting their applicability across diverse scenarios. In this paper, we propose AnchorHOI, a novel framework that thoroughly exploits hybrid priors by incorporating video diffusion models beyond image diffusion models, advancing 4D HOI generation. Nevertheless, directly optimizing high-dimensional 4D HOI with such priors remains challenging, particularly for human pose and compositional motion. To address this challenge, AnchorHOI introduces an anchor-based prior distillation strategy, which constructs interaction-aware anchors and then leverages them to guide generation in a tractable two-step process. Specifically, two tailored anchors are designed for 4D HOI generation: anchor Neural Radiance Fields (NeRFs) for expressive interaction composition, and anchor keypoints for realistic motion synthesis. Extensive experiments demonstrate that AnchorHOI outperforms previous methods with superior diversity and generalization.

4D生成扩散模型人-物交互零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。