arXiv:2503.18349cs.CV2025-03被引 9

用视觉语言模型自动设计动作策略,实现自然的人物-物体交互生成

Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy

  • 基于VLM自动生成任务目标与奖励函数,无需人工调参
  • 在数千条长时序交互数据上验证,生成动作更自然、多样
  • 适合动画、仿真和机器人领域,支持静态/动态/关节物体

人-物交互(HOI)合成在动画、仿真和机器人中至关重要。现有方法或依赖昂贵的动作捕捉数据,或需人工设计奖励函数,限制了可扩展性和泛化能力。本文提出首个统一的物理驱动HOI框架,利用视觉语言模型(VLMs)实现对多种物体类型(包括静态、动态和铰接物体)的长时序交互。我们引入VLM-Guided相对运动动力学(RMD),一种细粒度时空二分图表示,可自动构建目标状态与强化学习奖励函数。通过编码人体与物体部件间的结构关系,RMD使VLM生成语义一致、交互感知的动作引导,无需手动奖励调优。为支持该方法,我们构建了Interplay数据集,包含数千条长时序静态与动态交互计划。大量实验表明,该框架在单任务与多任务复杂场景下均优于现有方法,能生成更自然、类人的动作。更多细节请见项目网页:https://vlm-rmd.github.io/

原文摘要 · Abstract (English)

Human-object interaction (HOI) synthesis is crucial for applications in animation, simulation, and robotics. However, existing approaches either rely on expensive motion capture data or require manual reward engineering, limiting their scalability and generalizability. In this work, we introduce the first unified physics-based HOI framework that leverages Vision-Language Models (VLMs) to enable long-horizon interactions with diverse object types, including static, dynamic, and articulated objects. We introduce VLM-Guided Relative Movement Dynamics (RMD), a fine-grained spatio-temporal bipartite representation that automatically constructs goal states and reward functions for reinforcement learning. By encoding structured relationships between human and object parts, RMD enables VLMs to generate semantically grounded, interaction-aware motion guidance without manual reward tuning. To support our methodology, we present Interplay, a novel dataset with thousands of long-horizon static and dynamic interaction plans. Extensive experiments demonstrate that our framework outperforms existing methods in synthesizing natural, human-like motions across both simple single-task and complex multi-task scenarios. For more details, please refer to our project webpage: https://vlm-rmd.github.io/.

人物交互视觉语言模型动作生成强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。