arXiv:2409.10090cs.CV2024-09被引 1

让图像生成自动规划物体运动,真实模拟物理互动。

InteractPro: A Unified Framework for Motion-Aware Image Composition

  • 用大模型智能分析场景并决定物体摆放策略。
  • 结合物理模拟与扩散模型,生成动态逼真的图像组合。
  • 无需训练即可实现可控、连贯的运动图像生成,适合影视动画设计。

我们提出 InteractPro,一个用于动态运动感知图像合成的统一框架。核心是 InteractPlan,一个基于大视觉语言模型(LVLM)的智能规划器,可进行场景分析与物体布局,确定最优合成策略以实现逼真运动效果。根据具体场景,InteractPlan 选择两个专用模块之一:InteractPhys 采用改进的材料点法(MPM)模拟,实现高保真且可控的物体-场景交互,能捕捉需真实物理建模的多样抽象事件;InteractMotion 则为基于预训练视频扩散模型的免训练方法。传统合成方法存在两大缺陷:需手动规划物体位置,且输出为静态图像。InteractPro 在规划器引导下融合模拟与扩散方法,克服上述局限,确保丰富运动感知的合成效果。大量定量与定性评估证明其在多样化场景中均能生成可控且连贯的图像组合。

原文摘要 · Abstract (English)

We introduce InteractPro, a comprehensive framework for dynamic motion-aware image composition. At its core is InteractPlan, an intelligent planner that leverages a Large Vision Language Model (LVLM) for scenario analysis and object placement, determining the optimal composition strategy to achieve realistic motion effects. Based on each scenario, InteractPlan selects between our two specialized modules: InteractPhys and InteractMotion. InteractPhys employs an enhanced Material Point Method (MPM)-based simulation to produce physically faithful and controllable object-scene interactions, capturing diverse and abstract events that require true physical modeling. InteractMotion, in contrast, is a training-free method based on pretrained video diffusion. Traditional composition approaches suffer from two major limitations: requiring manual planning for object placement and generating static, motionless outputs. By unifying simulation-based and diffusion-based methods under planner guidance, InteractPro overcomes these challenges, ensuring richly motion-aware compositions. Extensive quantitative and qualitative evaluations demonstrate InteractPro's effectiveness in producing controllable, and coherent compositions across varied scenarios.

图像合成物理模拟视频生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。