arXiv:2608.24119cs.CVcs.AI2026-08

让图像编辑理解物理规律,能自动推理并适配场景的视觉编辑框架

TransPhy: Visual In-Context Learning for Physically Grounded Image Editing

论文配图:TransPhy: Visual In-Context Learning for Physically Grounded Image Editing
图 1 · 摘自论文原文
  • 分解物理规则推断与对齐渲染,通过专家混合机制自适应生成
  • 在74条物理规则、5240对图像上实现更高规则遵循度和泛化能力
  • 适合需要真实物理交互的图像编辑场景,如设计、影视特效

视觉示范为难以用文字详尽描述的图像变换提供了自然接口。然而,现有视觉上下文学习(VICL)方法主要关注外观关系迁移,对依赖材料属性、几何结构、物体交互和环境条件的物理基础变换支持有限。给定源-目标示例对和查询图像,物理基础VICL需模型推断演示变换、适配查询特定场景上下文,并保留无关规则内容。我们提出PhysVICL-74,包含74条物理基础变换规则和5240对源-目标图像,形成近75,000个训练与评估上下文。其基准划分独立评估新实例迁移与未见规则泛化。我们进一步提出TransPhy框架,将物理基础VICL分解为物理规则诱导与过渡对齐渲染。TransPhy先预测演示规则和显式的查询特定目标状态描述,再通过逐标记专家混合机制合成目标图像,专家路由由局部过渡线索引导。实验表明,TransPhy在物理规则遵循度、查询一致性及未见规则泛化方面优于现有视觉上下文编辑方法。

原文摘要 · Abstract (English)

Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation transfer and provide limited support for physically grounded transformations, whose outcomes depend on material properties, geometry, object interactions, and environmental conditions. Given a source--target exemplar pair and a query image, physically grounded VICL requires a model to infer the demonstrated transformation, adapt its effects to the query-specific scene context, and preserve rule-irrelevant content. We introduce PhysVICL-74, comprising 74 physically grounded transformation rules and 5,240 source--target image pairs that form nearly 75K training and evaluation contexts. Its benchmark split separately evaluates novel-instance transfer and unseen-rule generalization. We further propose TransPhy, a framework that decomposes physically grounded VICL into physical-rule induction and transition-aligned rendering. TransPhy first predicts the demonstrated rule and an explicit query-specific target-state description, and then synthesizes the target image through token-wise mixture-of-experts adaptation, with expert routing guided by localized transition cues. Experiments show that TransPhy improves physical-rule adherence, query consistency, and unseen-rule generalization over existing visual in-context editing methods.

图像编辑物理建模视觉学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。