arXiv:2602.01881cs.CV2026-02

用分层代理嵌入实现可控图像编辑,支持精细操作与实时动画。

ProxyImg: Towards Highly-Controllable Image Representation via Hierarchical Disentangled Proxy Embedding

  • 分层代理几何结构解耦语义、几何与纹理,参数独立可调。
  • 在ImageNet等数据集上以更少参数达顶尖重建质量,支持无生成模型的背景补全。
  • 适合需要精细控制、实时动画的视觉内容创作人群。

现有图像表示方法,如栅格图像、高斯原型和隐式表征,或存在冗余导致需大量手动编辑,或缺乏潜在变量到语义实例/部分的直接映射,难以实现细粒度操控。为此,我们提出一种基于分层代理的参数化图像表示,将语义、几何和纹理属性解耦至独立可调的参数空间。通过语义感知的图像分解,利用自适应贝塞尔拟合与迭代区域分割及网格化构建分层代理几何结构,并在几何感知的分布式代理节点中嵌入多尺度隐式纹理参数,实现像素域连续高保真重建,支持与实例或部件无关的语义编辑。此外,引入局部自适应特征索引机制保障空间纹理连贯性,实现无需生成模型的高质量背景补全。在ImageNet、OIR-Bench和HumanEdit等多个重建与编辑基准上实验表明,该方法以显著更少参数达到顶尖渲染保真度,且支持直观、交互式、物理合理的操作。结合位置动力学(Position-Based Dynamics),框架可实现轻量隐式渲染下的实时物理驱动动画,时间一致性与视觉真实感优于生成式方法。

原文摘要 · Abstract (English)

Prevailing image representation methods, including explicit representations such as raster images and Gaussian primitives, as well as implicit representations such as latent images, either suffer from representation redundancy that leads to heavy manual editing effort, or lack a direct mapping from latent variables to semantic instances or parts, making fine-grained manipulation difficult. These limitations hinder efficient and controllable image and video editing. To address these issues, we propose a hierarchical proxy-based parametric image representation that disentangles semantic, geometric, and textural attributes into independent and manipulable parameter spaces. Based on a semantic-aware decomposition of the input image, our representation constructs hierarchical proxy geometries through adaptive Bezier fitting and iterative internal region subdivision and meshing. Multi-scale implicit texture parameters are embedded into the resulting geometry-aware distributed proxy nodes, enabling continuous high-fidelity reconstruction in the pixel domain and instance- or part-independent semantic editing. In addition, we introduce a locality-adaptive feature indexing mechanism to ensure spatial texture coherence, which further supports high-quality background completion without relying on generative models. Extensive experiments on image reconstruction and editing benchmarks, including ImageNet, OIR-Bench, and HumanEdit, demonstrate that our method achieves state-of-the-art rendering fidelity with significantly fewer parameters, while enabling intuitive, interactive, and physically plausible manipulation. Moreover, by integrating proxy nodes with Position-Based Dynamics, our framework supports real-time physics-driven animation using lightweight implicit rendering, achieving superior temporal consistency and visual realism compared with generative approaches.

图像表示可控编辑代理嵌入物理动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。