arXiv:2601.10103cs.CVcs.AI2026-01被引 9

FlowAct-R1实现低延迟实时交互式人形视频生成,支持自然动作切换。

FlowAct-R1: Towards Interactive Humanoid Video Generation

  • 基于MMDiT架构与分块扩散强迫策略,支持任意时长视频流式生成
  • 25fps稳定输出,首帧延迟仅1.5秒,长期时间一致性良好
  • 适合需要高实时性与精细全身控制的交互式虚拟角色应用

交互式人形视频生成旨在合成能通过连续响应视频与人类互动的逼真视觉代理。尽管视频合成技术取得进展,现有方法常在高保真度与实时交互需求间面临权衡。本文提出FlowAct-R1框架,专为实时交互式人形视频生成设计。基于MMDiT架构,该框架支持任意时长视频的流式合成,并保持低延迟响应。引入分块扩散强迫策略及新型自强迫变体,缓解误差累积,保障长时间序列的一致性。通过高效蒸馏与系统级优化,框架在480p分辨率下实现稳定25fps输出,首帧时间(TTFF)约1.5秒。所提方法提供整体且细粒度的全身控制,使代理可在交互场景中自然切换多种行为状态。实验表明,FlowAct-R1在行为生动性与感知真实感上表现卓越,同时对不同角色风格具备强泛化能力。

原文摘要 · Abstract (English)

Interactive humanoid video generation aims to synthesize lifelike visual agents that can engage with humans through continuous and responsive video. Despite recent advances in video synthesis, existing methods often grapple with the trade-off between high-fidelity synthesis and real-time interaction requirements. In this paper, we propose FlowAct-R1, a framework specifically designed for real-time interactive humanoid video generation. Built upon a MMDiT architecture, FlowAct-R1 enables the streaming synthesis of video with arbitrary durations while maintaining low-latency responsiveness. We introduce a chunkwise diffusion forcing strategy, complemented by a novel self-forcing variant, to alleviate error accumulation and ensure long-term temporal consistency during continuous interaction. By leveraging efficient distillation and system-level optimizations, our framework achieves a stable 25fps at 480p resolution with a time-to-first-frame (TTFF) of only around 1.5 seconds. The proposed method provides holistic and fine-grained full-body control, enabling the agent to transition naturally between diverse behavioral states in interactive scenarios. Experimental results demonstrate that FlowAct-R1 achieves exceptional behavioral vividness and perceptual realism, while maintaining robust generalization across diverse character styles.

人形生成实时视频扩散模型交互控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。