arXiv:2607.20174cs.CVcs.AI2026-07

StreamHOI实现低延迟长时序人物交互视频生成,兼顾流畅性与细节。

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation

论文配图:StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation
图 1 · 摘自论文原文
  • 设计动态记忆架构,按不同Transformer块特性分配历史记忆
  • 生成速度达17.6帧/秒,首段延迟仅0.75秒,支持实时交互
  • 适合需要高保真人物交互的流式视频生成应用

现有交互视频生成方法多局限于离线短视频生成,难以满足实时交互需求。本文提出StreamHOI,一种面向长时序人-物交互视频生成的低延迟流式框架。不同于将复杂条件模型直接转为流系统,我们研究图像到视频生成器如何组织历史记忆以在有限延迟下保持交互一致性。发现标准局部记忆设计存在权衡,不同Transformer模块对交互区域与背景区域的记忆偏好各异。为此,StreamHOI通过离线块级感知分析,实施偏置引导的记忆专化训练,适配模块特异性记忆布局,并引入记忆距离缩放模块强化对早期交互状态的远距离访问能力。大量实验对比显示,StreamHOI在交互合理性、物体保真度、人物质量与效率上均表现优异,达到17.6 FPS,首段延迟0.75秒。

原文摘要 · Abstract (English)

Existing human--object interaction (HOI) video generation methods are largely limited to offline short-video generation with complex driving conditions, making them unsuitable for real-time interactive applications. We present \emph{StreamHOI}, a low-latency streaming framework for long-duration HOI video generation. Instead of converting heavily conditioned HOI pipelines into streaming systems, we study how an image-to-video streaming generator should organize historical memory to preserve interactions under bounded latency. We find that the standard sink-local memory design faces a trade-off in streaming HOI generation, and different transformer blocks show different historical-memory preferences for HOI regions and surrounding regions. To match memory composition with block behavior, StreamHOI performs offline HOI-aware block profiling and applies bias-guided memory-specialized training to adapt the generator to block-specific memory layouts. We further introduce a memory distance scaling module to strengthen long-range access to early interaction states. Extensive comparisons with both long-video baselines and recent HOI generation methods demonstrate that StreamHOI achieves strong interaction plausibility, object fidelity, human quality and efficiency, reaching 17.6 FPS with 0.75s first-chunk latency.

视频生成交互建模流式处理注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。