arXiv:2608.12232cs.CV2026-08

无需3D重建即可实现物体几何感知的视频缩放,保持形状合理与背景一致。

ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference

论文配图:ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference
图 1 · 摘自论文原文
  • 分两阶段训练:先学前后景合成,再引入物体中心3D形变指导缩放。
  • 在真实视频上生成高质量结果,无需配对的真实缩放数据。
  • 速度快、效果优,适合需要高效视频编辑的开发者和创作者。

几何感知的视频物体缩放旨在沿物体中心轴非均匀地调整物体大小,同时保持几何合理性、时间连贯性和背景一致性。现有文本引导方法主要在2D图像平面操作,深度引导方法控制粗糙,基于网格的方法需昂贵的3D重建。本文提出一种渐进式两阶段训练框架,将几何感知前景变换与背景保持、真实视频合成解耦,推理时无需网格-像素对齐或显式3D重建。两个阶段均从真实视频构造几何扰动伪源,原完整视频作为重建目标。第一阶段使用平面变换学习鲁棒的前景-背景组合,第二阶段引入物体中心3D形变引导实现几何感知缩放。该伪源重建范式可在无成对真实缩放目标下实现真实视频生成。我们构建了互补的配对几何与真实背景基准,并在真实场景视频上进一步评估。大量实验表明,在几何一致性、前景保真度和背景保持方面优于现有方法,且推理更快更实用。

原文摘要 · Abstract (English)

Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, temporal coherence, and background consistency. Existing text-guided methods mainly operate in the 2D image plane, while depth-guided approaches provide coarse control and mesh-based methods require costly 3D reconstruction. We present a progressive two-stage training framework that decouples geometry-aware foreground transformation from background preservation and realistic video composition, without mesh-pixel alignment and explicit 3D reconstruction at inference. In both stages, geometrically perturbed pseudo-sources are constructed from real videos, while the original complete videos are retained as reconstruction targets. The first stage uses planar transformations to learn robust foreground-background composition, whereas the second introduces object-centric 3D deformation guidance for geometry-aware scaling. This pseudo-source reconstruction formulation enables real-video synthesis without paired real-world scaling targets. We construct complementary paired-geometry and real-background benchmarks and further evaluate on in-the-wild videos. Extensive experiments demonstrate superior geometric consistency, foreground fidelity, and background preservation, together with faster and more practical inference than methods requiring explicit 3D reconstruction.

视频编辑几何感知高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。