PINO让任意规模人群互动生成更真实可控,无需训练即可实现精细动作设计。
PINO: Person-Interaction Noise Optimization for Long-Duration and Customizable Motion Generation of Arbitrary-Sized Groups
- 将群体互动拆解为两两关系,用预训练双人模型逐步组合生成
- 通过物理惩罚避免角色穿模重叠,确保动作符合现实力学
- 支持用户自定义角色朝向、速度和空间关系,适用于游戏与机器人
多人角色群体交互的生成因群体规模扩大而日益复杂。现有条件扩散模型虽能逐个生成角色动作,但依赖单一共享提示,控制力弱且交互过于简单。本文提出无需训练的Person-Interaction Noise Optimization(PINO)框架,可生成任意规模群体的逼真、可定制交互。PINO将复杂群体互动分解为语义相关的成对交互,利用预训练的双人交互扩散模型逐步构建整体动作。为保证物理合理性,避免角色穿模或重叠等常见伪影,PINO在噪声优化中引入基于物理的惩罚机制。该方法可在不增加训练成本的前提下,精确控制角色朝向、速度与空间关系。大量实验表明,PINO生成的动作视觉真实、物理一致且适应性强,适用于动画、游戏及机器人等多种应用场景。
原文摘要 · Abstract (English)
Generating realistic group interactions involving multiple characters remains challenging due to increasing complexity as group size expands. While existing conditional diffusion models incrementally generate motions by conditioning on previously generated characters, they rely on single shared prompts, limiting nuanced control and leading to overly simplified interactions. In this paper, we introduce Person-Interaction Noise Optimization (PINO), a novel, training-free framework designed for generating realistic and customizable interactions among groups of arbitrary size. PINO decomposes complex group interactions into semantically relevant pairwise interactions, and leverages pretrained two-person interaction diffusion models to incrementally compose group interactions. To ensure physical plausibility and avoid common artifacts such as overlapping or penetration between characters, PINO employs physics-based penalties during noise optimization. This approach allows precise user control over character orientation, speed, and spatial relationships without additional training. Comprehensive evaluations demonstrate that PINO generates visually realistic, physically coherent, and adaptable multi-person interactions suitable for diverse animation, gaming, and robotics applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。