实时生成人与环境互动的反应动作,支持低延迟高保真输出。
ReMoGen: Real-time Human Interaction-to-Reaction Generation via Modular Learning from Diverse Data
- 模块化学习框架,分步适配多源交互数据
- 帧级细化提升响应速度与动作连贯性
- 适合虚拟角色、人机协作等实时交互场景
真实世界中人类行为具有高度互动性,个体动作受周围对象与场景影响。本文针对实时交互到反应的动作生成任务,从多人动作、场景几何及语义输入等多源动态线索中生成主体未来动作。该任务面临两大挑战:一是交互数据分散于单人、人-人、人-场景等异构领域且数据稀疏;二是需在连续在线交互中实现低延迟、高保真响应。为此提出ReMoGen(Reaction Motion Generation)框架,利用大规模单人动作数据学习通用运动先验,并通过独立训练的Meta-Interaction模块适配目标交互领域,实现数据稀缺下的鲁棒泛化。为支持实时响应,采用分段生成结合轻量级帧级修正模块,在每帧融合新观测线索,提升响应速度与时间一致性,无需全序列重推断。跨人-人、人-场景及多模态交互设置的大量实验表明,ReMoGen生成动作质量高、连贯性强、响应快,且在多样交互场景中具有良好泛化能力。
原文摘要 · Abstract (English)
Human behaviors in real-world environments are inherently interactive, with an individual's motion shaped by surrounding agents and the scene. Such capabilities are essential for applications in virtual avatars, interactive animation, and human-robot collaboration. We target real-time human interaction-to-reaction generation, which generates the ego's future motion from dynamic multi-source cues, including others' actions, scene geometry, and optional high-level semantic inputs. This task is fundamentally challenging due to (i) limited and fragmented interaction data distributed across heterogeneous single-person, human-human, and human-scene domains, and (ii) the need to produce low-latency yet high-fidelity motion responses during continuous online interaction. To address these challenges, we propose ReMoGen (Reaction Motion Generation), a modular learning framework for real-time interaction-to-reaction generation. ReMoGen leverages a universal motion prior learned from large-scale single-person motion datasets and adapts it to target interaction domains through independently trained Meta-Interaction modules, enabling robust generalization under data-scarce and heterogeneous supervision. To support responsive online interaction, ReMoGen performs segment-level generation together with a lightweight Frame-wise Segment Refinement module that incorporates newly observed cues at the frame level, improving both responsiveness and temporal coherence without expensive full-sequence inference. Extensive experiments across human-human, human-scene, and mixed-modality interaction settings show that ReMoGen produces high-quality, coherent, and responsive reactions, while generalizing effectively across diverse interaction scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。