通过几何与语义通道分离相机和主体运动,实现可控视频生成的精确控制。
OrthoMotion:Disentangling Camera and Subject Motion via Geometry Semantics Orthogonal Attention

- 用旋转位置编码和门控值注入分别处理相机与主体运动
- 在多个模型上实现超过2.4倍的交叉干扰降低,精度领先
- 首次从构造上保证解耦,适合需要精准运动控制的应用
可控视频生成要求独立控制相机和主体运动,但2D条件下的运动流形共享相同的逆深度(1/Z)缩放,仅凭图像证据无法分离。我们证明这种纠缠是表示性的,而非架构问题——2D下相机与主体的分离是一个不可识别的逆问题,因此将解耦重构为算子设计问题。OrthoMotion在注意力算子层面解决此问题:将相机运动引导至几何通道,通过旋转位置嵌入(RoPE)相位实现保范变换;将主体运动引导至语义通道,通过交叉注意力中的门控值注入。由于这两个子算子代数互补(旋转与平移的仿射作用),轻量级正则化可保证其响应子空间正交,从而消除两者的相互干扰。据我们所知,OrthoMotion是首个通过构造实现解耦的方法,同时达到当前最优的相机与主体运动准确率,交叉干扰降低超2.4倍,且不损失保真度,可泛化至多种骨干网络。
原文摘要 · Abstract (English)
Controllable video generation demands independent command of the camera and the subject, yet 2D conditioning entangles them: camera- and object-induced optical flow share the same inverse-depth (1/Z) scaling and cannot be separated from image evidence alone. We first prove that this entanglement is representational, not architectural -- the 2D camera/object split is a non-identifiable inverse problem -- and therefore reframe decoupling as a question of operator design. We resolve it at the level of the attention operator. OrthoMotion routes camera motion into a geometric channel, a norm-preserving rotation of the rotary position embedding (RoPE) phase, and subject motion into a semantic channel, a gated value injection in cross-attention. Because these sub-operators are algebraically complementary -- a rotation versus a translation of the affine action on tokens -- a lightweight decoupling regularizer provably drives their response subspaces to orthogonality, so the two controls stop interfering. To our knowledge OrthoMotion is the first method to guarantee disentanglement by construction rather than hope for it to emerge. It attains state-of-the-art camera and subject accuracy at once while minimizing cross-talk, which we quantify with a new Cross-Talk Error (CTE) metric, cutting cross-talk by more than 2.4x with no loss in fidelity and generalizing across backbones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。