让视频物体插入更真实,基于环境感知的推理能力
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
- 用多模态大模型分析场景,判断物体放置的物理合理性
- 生成语义与空间双重引导,确保插入位置符合物理规律
- 适合需要高真实感视频编辑的研究者与创作者
视频物体插入是视频编辑的基础任务,但现有扩散方法常产生视觉合理却物理不一致的结果。本文提出 Place-it-R1,一种端到端框架,通过环境感知的多模态大模型(MLLM)推理实现物理合理的视频物体插入。不同于将推理视为通用文本提示,Place-it-R1利用 MLLM 分析目标环境,推断物体-场景交互,并确定插入在物理上有效的区域。由此生成的推理转化为两种互补的扩散引导:描述预期物理交互的语义引导,以及每帧中有效插入区域的空间引导。为进一步对齐生成结果与局部物理真实性,引入空间直接偏好优化(Spatial Direct Preference Optimization),利用 MLLM 对生成候选进行排序,并设计区域感知偏好目标,将物理违规惩罚显式定位至插入物体区域。Place-it-R1 还提供灵活的标准模式,可在环境适应性与场景保真度间权衡。大量实验表明,该方法生成的插入结果比当前最优方法更具物理一致性与视觉自然性,且在性能上可媲美商业系统。
原文摘要 · Abstract (English)
Video object insertion is fundamental to video editing, yet existing diffusion methods often produce visually plausible but physically inconsistent results. We present Place-it-R1, an end-to-end framework for physically plausible video object insertion driven by environment-aware MLLM reasoning. Rather than treating reasoning as a generic text prompt, Place-it-R1 uses the MLLM to analyze the target environment, infer object-scene interactions, and determine where an insertion is physically valid. The resulting reasoning is translated into two complementary forms of guidance for video diffusion: semantic guidance that describes the intended physical interaction and spatial guidance that provides a valid insertion region in each frame. To further align generation with local physical realism, we introduce Spatial Direct Preference Optimization, which leverages an MLLM to rank generated candidates, and introduces a region-aware preference objective that explicitly localizes physical-violation penalties to the inserted-object region. Place-it-R1 further offers flexible and standard modes to trade off environment adaptation and scene preservation. Extensive experiments show that Place-it-R1 produces more physically coherent and visually natural insertions than state-of-the-art methods and achieves competitive results against commercial systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。