arXiv:2609.06628cs.CV2026-09

让视频物体位置和大小可精准编辑,且不因换外观而乱跑。

GeoCo-SAVi: Geometry-Consistent Slot Attention for Explicitly Editable Object Representations

论文配图:GeoCo-SAVi: Geometry-Consistent Slot Attention for Explicitly Editable Object Representations
图 1 · 摘自论文原文
  • 用空间等变解码器让位置缩放指令直接控制物体渲染位置与大小。
  • 在Obj3D上位置误差降低80%以上,换外观时尺寸变化减少50%以上。
  • 适合需要精确控制物体布局的视频生成与编辑任务。

物体中心视频模型用槽位表示场景,但外观变化会导致几何含义不一致。在不变槽注意力(ISA)中,显式的位置与尺度可能与解码后的中心和范围不符;编辑时会出现意外运动或缩放,替换外观还会导致几何偏移。GeoCo-SAVi增强几何权威性与语义一致性:其空间等变的物体级解码器使位置和尺度成为有效指令,改变它们可准确移动或缩放渲染对象。真实位置对齐将位置绑定至解码中心,归一化注意力重叠避免重复分配。外观移植使几何语义对齐,接收者几何决定布局,捐赠者外观提供形状。时间初始化器跨帧传播校准槽位。在Obj3D上,GeoCo-SAVi达到与ISA相当的重建效果,位置中心误差下降超80%,外观引起的尺寸波动减少超50%,并实现预期的平移与缩放响应。在250 MOVi-C视频上,也优于两个同协议基线,在重建、实例分组与固定身份几何控制方面表现更优。GeoCo-SAVi将显式几何转化为组合控制,使位置与尺度更可读、可编辑。

原文摘要 · Abstract (English)

Object-centric video models represent scenes with slots, yet exposed geometry can vary in meaning with appearance. In Invariant Slot Attention (ISA), explicit position and scale can disagree with the decoded center and extent; edits can yield unexpected motion or resizing, and replacing appearance can shift geometry. GeoCo-SAVi promotes geometric authority and semantic alignment. Its spatially equivariant, object-wise decoder makes position and scale effective commands: changing them moves or resizes the rendered support. Factual position alignment ties position to the decoded center, and normalized attention overlap discourages duplicate allocation. Appearance transplantation aligns geometry semantics across objects, so recipient geometry governs layout while donor appearance supplies shape. A temporal initializer propagates calibrated slots across frames. On Obj3D, GeoCo-SAVi matches ISA reconstruction, reduces p-centroid error by over 80%, and cuts appearance-induced size variation by over 50% while producing the expected translation and scale responses. On 250 MOVi-C videos, it also improves reconstruction, instance grouping, and fixed-identity geometry control over two same-protocol references. GeoCo-SAVi transforms explicit geometry into compositional control, making both position and scale more readable and editable.

视频生成槽注意力几何控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。