arXiv:2604.19636cs.CV2026-04被引 2

让人体与物体交互视频更真实,避免手穿模等问题。

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

论文配图:CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
图 1 · 摘自论文原文
  • 用空间引导的专家路由提升手脸等部位结构精度
  • 双流训练注入交互几何先验,生成无穿模动作
  • 适合电商、虚拟营销等需要逼真交互视频的场景

人体-物体交互(HOI)视频合成在电商、数字广告和虚拟营销中具有广泛应用价值。然而,现有扩散模型虽能生成逼真图像,却常出现(i)手部、面部等敏感区域结构不稳,以及(ii)物理上不合理的接触(如手物穿模)问题。本文提出CoInteract,一种端到端框架,基于人物参考图、产品参考图、文本提示和语音音频生成HOI视频。该框架嵌入于扩散Transformer(DiT)主干网络,引入两项互补设计:首先,提出人类感知的专家混合(MoE),通过空间监督路由机制将令牌分配至轻量级区域专用专家,以极小参数开销提升细粒度结构保真度;其次,提出空间结构化联合生成,采用双流训练范式,同时建模RGB外观流与辅助的HOI结构流,以注入交互几何先验。训练时,结构流关注RGB令牌并监督共享主干权重;推理时,移除结构分支实现零开销的RGB生成。实验表明,CoInteract在结构稳定性、逻辑一致性和交互真实性方面显著优于现有方法。

原文摘要 · Abstract (English)

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capability, still frequently fail on (i) the structural stability of sensitive regions such as hands and faces and (ii) physically plausible contact (e.g., avoiding hand--object interpenetration). We present CoInteract, an end-to-end framework for HOI video synthesis conditioned on a person reference image, a product reference image, text prompts, and speech audio. CoInteract introduces two complementary designs embedded into a Diffusion Transformer (DiT) backbone. First, we propose a Human-Aware Mixture-of-Experts (MoE) that routes tokens to lightweight, region-specialized experts via spatially supervised routing, improving fine-grained structural fidelity with minimal parameter overhead. Second, we propose Spatially-Structured Co-Generation, a dual-stream training paradigm that jointly models an RGB appearance stream and an auxiliary HOI structure stream to inject interaction geometry priors. During training, the HOI stream attends to RGB tokens and its supervision regularizes shared backbone weights; at inference, the HOI branch is removed for zero-overhead RGB generation. Experimental results demonstrate that CoInteract significantly outperforms existing methods in structural stability, logical consistency, and interaction realism.

视频生成扩散模型人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。