arXiv:2602.21778cs.CV2026-02被引 1

让图像编辑模拟物理动态,生成更真实的材质变形与光路变化效果。

From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors

  • 将编辑任务转为物理状态预测,用时序轨迹建模动态过程。
  • 在38K条物理过渡数据上训练,提升复杂动态编辑的逼真度。
  • 适合需要真实物理模拟的图像编辑场景,如影视特效与工业设计。

基于指令的图像编辑在语义对齐方面已取得显著进展,但当前先进模型在涉及折射、材料变形等复杂因果动态的编辑任务中常产生不合理的物理结果。我们指出,这源于主流范式将编辑视为图像对之间的离散映射,仅提供边界条件而未指定过渡动态。为此,我们将物理感知编辑重新定义为可预测的物理状态转移,并构建了包含38,000条跨五个物理领域的过渡轨迹的大规模视频数据集PhysicTran38K,通过两阶段过滤与约束感知标注流程生成。在此监督基础上,提出PhysicEdit端到端框架,融合冻结的Qwen2.5-VL进行物理推理与可学习的时序查询,为扩散模型提供自适应视觉引导。实验表明,PhysicEdit在物理真实性上较Qwen-Image-Edit提升5.9%,在知识驱动编辑上提升10.1%,成为开源方法新标杆,同时媲美领先专有模型。

原文摘要 · Abstract (English)

Instruction-based image editing has achieved remarkable success in semantic alignment, yet state-of-the-art models frequently fail to render physically plausible results when editing involves complex causal dynamics, such as refraction or material deformation. We attribute this limitation to the dominant paradigm that treats editing as a discrete mapping between image pairs, which provides only boundary conditions and leaves transition dynamics underspecified. To address this, we reformulate physics-aware editing as predictive physical state transitions and introduce PhysicTran38K, a large-scale video-based dataset comprising 38K transition trajectories across five physical domains, constructed via a two-stage filtering and constraint-aware annotation pipeline. Building on this supervision, we propose PhysicEdit, an end-to-end framework equipped with a textual-visual dual-thinking mechanism. It combines a frozen Qwen2.5-VL for physically grounded reasoning with learnable transition queries that provide timestep-adaptive visual guidance to a diffusion backbone. Experiments show that PhysicEdit improves over Qwen-Image-Edit by 5.9% in physical realism and 10.1% in knowledge-grounded editing, setting a new state-of-the-art for open-source methods, while remaining competitive with leading proprietary models.

图像编辑物理模拟扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。