用多智能体推理实现文本驱动的交互视频生成,无需显式动作控制。
AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment

- 通过多智能体协同推理规划感知、交互与运动
- 隐式对齐文本与动作,生成更自然的交互视频
- 适合需要复杂交互控制的研究者与开发者
视频扩散模型的进展推动了人-物交互(HOI)视频生成的发展,该任务需在单主体动画之外实现精细的交互逻辑控制。然而,现有方法严重依赖显式动作控制,限制了跨物体和交互类型的可扩展性与泛化能力。本文提出AgentHOI,一种遵循‘先思考后生成’框架的文本驱动HOI视频生成方法,通过多智能体推理在高层文本意图与物理执行之间建立桥梁,涵盖感知、交互与运动规划。基于生成的交互计划,进一步强化文本驱动的动作理解,引入隐式文本-动作对齐策略,将文本到动作的先验知识融入视频扩散模型,使推理阶段无需显式动作输入即可实现稳健的HOI合成。实验表明,AgentHOI在穿着、骑乘等复杂物中心场景中显著提升交互自然性、物体外观保真度及对复杂文本指令的遵循能力。代码已开源:https://github.com/bone-11/agenthoi。
原文摘要 · Abstract (English)
Recent advances in video diffusion models have spurred interest in human-object interaction (HOI) video generation, which demands fine-grained control over interaction logic beyond single-subject animation. However, existing HOI methods rely heavily on explicit motion control, limiting scalability and generalization across diverse objects and interactions. In this study, we propose AgentHOI, a text-driven HOI video generation following a thinking-before-generation framework that bridges the gap between high-level textual intent and physical execution through multi-agent reasoning over perception, interaction, and motion planning. Building upon the generated interaction plans, we further strengthen text-driven motion understanding. We introduce an implicit text-motion alignment strategy that distills text-to-motion priors into the video diffusion model, enabling robust HOI synthesis without explicit motion inputs at inference. Experiments show that AgentHOI significantly improves interaction naturalness, object appearance preservation, and adherence to complex textual instructions across challenging object-centric scenarios such as wearing and riding. The code is available at https://github.com/bone-11/agenthoi.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。