统一建模人与人、物、场景的交互,生成自然连贯动作。
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts

- 用统一体积场表示多种交互实体,实现跨场景推理。
- 在三类交互任务中表现优异,能泛化到未见组合。
- 适合复杂环境下的动作生成研究与应用。
我们提出Uni-Inter,一个统一框架,支持人-人、人-物、人-场景等多种交互情境的运动生成。相比依赖任务特定设计且泛化能力有限的现有方法,Uni-Inter引入统一交互体(UIV),将异构交互实体编码为共享空间场,实现一致的关联推理与复合交互建模。运动生成被形式化为对UIV的关节级概率预测,使模型能够捕捉精细的空间依赖关系,生成连贯且上下文感知的行为。在三个代表性交互任务上的实验表明,Uni-Inter性能具有竞争力,并能良好泛化至新的实体组合。结果表明,统一建模复合交互为复杂环境中可扩展的动作合成提供了有前景的方向。
原文摘要 · Abstract (English)
We present Uni-Inter, a unified framework for human motion generation that supports a wide range of interaction scenarios: including human-human, human-object, and human-scene-within a single, task-agnostic architecture. In contrast to existing methods that rely on task-specific designs and exhibit limited generalization, Uni-Inter introduces the Unified Interactive Volume (UIV), a volumetric representation that encodes heterogeneous interactive entities into a shared spatial field. This enables consistent relational reasoning and compound interaction modeling. Motion generation is formulated as joint-wise probabilistic prediction over the UIV, allowing the model to capture fine-grained spatial dependencies and produce coherent, context-aware behaviors. Experiments across three representative interaction tasks demonstrate that Uni-Inter achieves competitive performance and generalizes well to novel combinations of entities. These results suggest that unified modeling of compound interactions offers a promising direction for scalable motion synthesis in complex environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。