用几何一致的物体输入生成连贯的人物-物体交互视频
ByteLoom: Weaving Geometry-Consistent Human-Object Interactions through Progressive Curriculum Learning
- 通过相对坐标图缓存机制实现物体6自由度精准控制
- 无需精细手部网格标注,仍能生成流畅交互动作
- 适合数字人、电商广告和机器人模仿学习场景
人物-物体交互(HOI)视频生成在数字人、电商、广告和机器人模仿学习中具有广阔应用前景。然而现有方法存在两大局限:一是缺乏有效机制注入物体多视角信息,导致跨视角一致性差;二是严重依赖精细的手部网格标注来建模交互遮挡。为此,我们提出ByteLoom,一种基于扩散变换器(DiT)的框架,通过简化的人体条件输入与3D物体输入,生成几何一致的HOI视频。我们首先提出RCM-cache机制,利用相对坐标图(RCM)作为通用表示,在保持物体几何一致性的同时精确控制6-DoF物体变换。为缓解HOI数据稀缺问题并利用现有数据集,我们设计了一种渐进式训练课程,提升模型能力并降低对手部网格的依赖。大量实验表明,该方法能忠实保留人物身份与物体多视角几何结构,同时实现平滑运动与物体操作。
原文摘要 · Abstract (English)
Human-object interaction (HOI) video generation has garnered increasing attention due to its promising applications in digital humans, e-commerce, advertising, and robotics imitation learning. However, existing methods face two critical limitations: (1) a lack of effective mechanisms to inject multi-view information of the object into the model, leading to poor cross-view consistency, and (2) heavy reliance on fine-grained hand mesh annotations for modeling interaction occlusions. To address these challenges, we introduce ByteLoom, a Diffusion Transformer (DiT)-based framework that generates realistic HOI videos with geometrically consistent object illustration, using simplified human conditioning and 3D object inputs. We first propose an RCM-cache mechanism that leverages Relative Coordinate Maps (RCM) as a universal representation to maintain object's geometry consistency and precisely control 6-DoF object transformations in the meantime. To compensate HOI dataset scarcity and leverage existing datasets, we further design a training curriculum that enhances model capabilities in a progressive style and relaxes the demand of hand mesh. Extensive experiments demonstrate that our method faithfully preserves human identity and the object's multi-view geometry, while maintaining smooth motion and object manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。