统一多模态条件生成人物物品交互视频,提升内容创作自动化水平
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

- 采用统一通道条件与门控局部上下文注意力,融合文本、图像、音频和姿态
- 在多种模态组合下实现当前最佳生成效果,支持工业级视频质量
- 构建专用评测基准HOIVG-Bench,推动该领域评估标准化
本文研究人-物交互视频生成(HOIVG),旨在基于文本、参考图像、音频和姿态等多模态条件生成高质量视频。该任务在电商演示、短视频制作和互动娱乐等实际应用中具有重要价值,但现有方法难以同时支持所有条件。我们提出OmniShow,一个端到端框架,可有效融合多模态输入并实现工业级性能。为解决可控性与质量的权衡,引入统一通道条件以高效注入图像与姿态信息,并设计门控局部上下文注意力确保音视频精准同步。针对数据稀缺问题,提出解耦-联合训练策略,通过多阶段训练与模型合并,充分利用异构子任务数据集。此外,为填补评估空白,建立专用于HOIVG的HOIVG-Bench基准。大量实验表明,OmniShow在多种多模态条件设置下均达到当前最优表现,为新兴的HOIVG任务确立了坚实标准。
原文摘要 · Abstract (English)
In this work, we study Human-Object Interaction Video Generation (HOIVG), which aims to synthesize high-quality human-object interaction videos conditioned on text, reference images, audio, and pose. This task holds significant practical value for automating content creation in real-world applications, such as e-commerce demonstrations, short video production, and interactive entertainment. However, existing approaches fail to accommodate all these requisite conditions. We present OmniShow, an end-to-end framework tailored for this practical yet challenging task, capable of harmonizing multimodal conditions and delivering industry-grade performance. To overcome the trade-off between controllability and quality, we introduce Unified Channel-wise Conditioning for efficient image and pose injection, and Gated Local-Context Attention to ensure precise audio-visual synchronization. To effectively address data scarcity, we develop a Decoupled-Then-Joint Training strategy that leverages a multi-stage training process with model merging to efficiently harness heterogeneous sub-task datasets. Furthermore, to fill the evaluation gap in this field, we establish HOIVG-Bench, a dedicated and comprehensive benchmark for HOIVG. Extensive experiments demonstrate that OmniShow achieves overall state-of-the-art performance across various multimodal conditioning settings, setting a solid standard for the emerging HOIVG task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。