用扩散Transformer生成高保真人品互动视频,精准保留人物与产品特征。
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
- 通过参考信息注入和掩码交叉注意力,同步保持人与产品的身份特征。
- 利用3D人体模板与产品框实现手势与物品的精确对齐,动作自然真实。
- 适合电商与数字营销场景,生成视频质量优于现有方法。
在电子商务与数字营销中,生成高保真的人品互动演示视频对有效展示产品至关重要。然而,现有框架往往无法同时保留人与产品的身份特征,且缺乏对人品空间关系的理解,导致表现不真实、交互不自然。为此,我们提出基于扩散Transformer(DiT)的框架,通过注入成对的人品参考信息,并采用额外的掩码交叉注意力机制,同步保留人类身份及产品特有细节(如标志、纹理)。我们使用3D身体网格模板和产品边界框提供精确运动引导,实现手部动作与产品位置的直观对齐。此外,通过结构化文本编码融入类别级语义,提升小角度旋转下帧间的3D一致性。在混合数据集上结合大量数据增强策略训练,本方法在保持人品身份完整性和生成真实演示动作方面优于当前最先进方法。
原文摘要 · Abstract (English)
In e-commerce and digital marketing, generating high-fidelity human-product demonstration videos is important for effective product presentation. However, most existing frameworks either fail to preserve the identities of both humans and products or lack an understanding of human-product spatial relationships, leading to unrealistic representations and unnatural interactions. To address these challenges, we propose a Diffusion Transformer (DiT)-based framework. Our method simultaneously preserves human identities and product-specific details, such as logos and textures, by injecting paired human-product reference information and utilizing an additional masked cross-attention mechanism. We employ a 3D body mesh template and product bounding boxes to provide precise motion guidance, enabling intuitive alignment of hand gestures with product placements. Additionally, structured text encoding is used to incorporate category-level semantics, enhancing 3D consistency during small rotational changes across frames. Trained on a hybrid dataset with extensive data augmentation strategies, our approach outperforms state-of-the-art techniques in maintaining the identity integrity of both humans and products and generating realistic demonstration motions. Project page: https://lizhenwangt.github.io/DreamActor-H1/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。