arXiv:2509.01232cs.CV2025-09被引 1

用多智能体图模型实现任意场景中长时序真人动作生成

FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework

  • 构建动态有向图上的多智能体系统,分阶段规划与生成动作
  • 在自建SceneBench上实现92%任务完成率,显著优于现有方法
  • 适合需要高真实感人体行为生成的研究者和开发者

人类-场景交互(HSI)旨在复杂环境中生成逼真的人类行为,但面临长时序、高层任务处理及未见场景泛化等挑战。为此,我们提出FantasyHSI,一种以视频生成为中心的无配对数据多智能体框架。通过将交互过程建模为动态有向图,构建包含场景导航代理(感知环境并规划路径)与规划代理(将长期目标分解为原子动作)的协作系统。关键创新在于引入评判代理,通过评估生成动作与计划路径的偏差,建立闭环反馈机制,动态纠正生成模型随机性导致的轨迹漂移,保障长期逻辑一致性。为提升动作物理真实性,采用直接偏好优化(DPO)训练动作生成器,显著减少肢体扭曲与脚部滑动等伪影。在自建SceneBench基准上的实验表明,FantasyHSI在泛化能力、长时序任务完成率和物理真实性方面均显著优于现有方法。

原文摘要 · Abstract (English)

Human-Scene Interaction (HSI) seeks to generate realistic human behaviors within complex environments, yet it faces significant challenges in handling long-horizon, high-level tasks and generalizing to unseen scenes. To address these limitations, we introduce FantasyHSI, a novel HSI framework centered on video generation and multi-agent systems that operates without paired data. We model the complex interaction process as a dynamic directed graph, upon which we build a collaborative multi-agent system. This system comprises a scene navigator agent for environmental perception and high-level path planning, and a planning agent that decomposes long-horizon goals into atomic actions. Critically, we introduce a critic agent that establishes a closed-loop feedback mechanism by evaluating the deviation between generated actions and the planned path. This allows for the dynamic correction of trajectory drifts caused by the stochasticity of the generative model, thereby ensuring long-term logical consistency. To enhance the physical realism of the generated motions, we leverage Direct Preference Optimization (DPO) to train the action generator, significantly reducing artifacts such as limb distortion and foot-sliding. Extensive experiments on our custom SceneBench benchmark demonstrate that FantasyHSI significantly outperforms existing methods in terms of generalization, long-horizon task completion, and physical realism. Ours project page: https://fantasy-amap.github.io/fantasy-hsi/

人体生成多智能体视频生成动作规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。