arXiv:2603.25734cs.CV2026-03被引 3

无需分类器,通过去噪速度自动实现人物交互的精准引导。

Unleashing Guidance Without Classifiers for Human-Object Interaction Animation

  • 利用去噪过程自身节奏生成引导信号,避免人工设计先验。
  • 在合成物体几何多样性训练下,接触语义对形状变化更鲁棒。
  • 相比传统方法,生成交互更真实,泛化能力更强,适合复杂场景动画。

生成逼真人-物交互(HOI)动画仍具挑战,需同时建模动态人体动作与多样的物体几何形态。以往基于扩散模型的方法常依赖手工设计的接触先验或人为设定的运动约束来提升接触质量。本文提出LIGHT,一种数据驱动的替代方案:引导信号源自去噪过程本身的节奏,减少对人工先验的依赖。基于扩散强迫思想,将表征分解为模态专属组件,并采用异步去噪调度分配个性化噪声水平。在此框架中,较清晰的组件通过交叉注意力引导较嘈杂的组件,实现无辅助分类器的引导。我们发现这种数据驱动的引导具有天然的接触感知能力,当训练中引入广泛合成的物体几何时,可进一步增强接触语义对几何多样性的不变性。大量实验表明,由去噪速度诱导的引导比传统无分类器引导更有效模拟接触先验的优势,同时实现更高的接触保真度、更真实的HOI生成以及更强的未见物体与任务泛化能力。

原文摘要 · Abstract (English)

Generating realistic human-object interaction (HOI) animations remains challenging because it requires jointly modeling dynamic human actions and diverse object geometries. Prior diffusion-based approaches often rely on hand-crafted contact priors or human-imposed kinematic constraints to improve contact quality. We propose LIGHT, a data-driven alternative in which guidance emerges from the denoising pace itself, reducing dependence on manually designed priors. Building on diffusion forcing, we factor the representation into modality-specific components and assign individualized noise levels with asynchronous denoising schedules. In this paradigm, cleaner components guide noisier ones through cross-attention, yielding guidance without auxiliary classifiers. We find that this data-driven guidance is inherently contact-aware, and can be enhanced when training is augmented with a broad spectrum of synthetic object geometries, encouraging invariance of contact semantics to geometric diversity. Extensive experiments show that pace-induced guidance more effectively mirrors the benefits of contact priors than conventional classifier-free guidance, while achieving higher contact fidelity, more realistic HOI generation, and stronger generalization to unseen objects and tasks.

人-物交互扩散模型生成动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。