用强化学习让物体按文字指令精准变形,支持旋转缩放等几何操作。
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
- 基于强化学习与扩散模型,通过文本指令控制物体几何变换。
- 在多个基准上实现高空间精度和场景一致性,优于现有方法。
- 无需成对数据,适合需要精确物体操作的视觉生成任务。
我们提出Talk2Move,一种基于强化学习的扩散框架,用于在场景中根据自然语言指令进行对象级几何变换。当前多模态生成系统难以通过自然语言操控物体的空间位置、方向或大小,因缺乏成对标注数据且受像素级优化限制。Talk2Move采用分组相对策略优化(GRPO),利用输入图像与轻量文本扰动生成多样化轨迹,避免昂贵的配对数据需求。空间奖励引导模型使几何变换与语言描述对齐,结合离策略步骤评估与主动采样机制,提升学习效率。此外,设计了以对象为中心的空间奖励,直接评估位移、旋转与缩放行为,实现可解释且连贯的变换。在精心构建的基准测试中,Talk2Move展现出精确、一致且语义忠实的物体变换能力,在空间准确性和场景一致性上均超越现有文本引导编辑方法。
原文摘要 · Abstract (English)
We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for multimodal generation systems. While existing text-based manipulation methods can adjust appearance or style, they struggle to perform object-level geometric transformations-such as translating, rotating, or resizing objects-due to scarce paired supervision and pixel-level optimization limits. Talk2Move employs Group Relative Policy Optimization (GRPO) to explore geometric actions through diverse rollouts generated from input images and lightweight textual variations, removing the need for costly paired data. A spatial reward guided model aligns geometric transformations with linguistic description, while off-policy step evaluation and active step sampling improve learning efficiency by focusing on informative transformation stages. Furthermore, we design object-centric spatial rewards that evaluate displacement, rotation, and scaling behaviors directly, enabling interpretable and coherent transformations. Experiments on curated benchmarks demonstrate that Talk2Move achieves precise, consistent, and semantically faithful object transformations, outperforming existing text-guided editing approaches in both spatial accuracy and scene coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。