arXiv:2605.12938cs.CVcs.AI2026-05被引 2

提出CRePE编码,让视频生成更稳定地应对广角和鱼眼镜头的相机控制。

CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation

论文配图:CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
图 1 · 摘自论文原文
  • 用深度感知的曲线射线分布表示图像位置,适配统一相机模型
  • 在多个几何与感知指标上优于基线,保持视频质量竞争力
  • 支持外部径向图控制,可实现场景结构生成与源视频运动迁移

相机条件视频生成需要在相机运动、镜头设置和场景结构变化下仍可靠的定位编码。现有注意力级相机编码要么仅提供射线信号,要么依赖针孔相机模型,难以适用于包含广角和鱼眼镜头的统一相机模型(UCM)。为此,我们提出曲线射线期望位置编码(CRePE),将每个图像标记表示为沿其源射线的深度感知位置分布,提供兼容统一相机模型的位置编码,捕捉广角与鱼眼相机引起的投影路径几何特性。CRePE通过添加至冻结视频DiTs的几何注意力适配器实现,注入逐标记的场景距离信息,并利用单目几何基础模型伪监督稳定优化。该设计提升了相机控制稳定性,改善多个几何感知与感知质量指标,同时保持视频质量竞争力。控制性位置编码消融实验表明,其综合平均排名优于基于RayRoPE的基线,验证了对多样化相机模型中投影路径整合的有效性。此外,通过将同一编码路径扩展至外部几何控制的径向混合强制(Radial MixForcing),CRePE还支持基于径向图的场景几何条件生成及源视频运动迁移。

原文摘要 · Abstract (English)

Camera-conditioned video generation requires positional encoding that remains reliable under changes in camera motion, lens configuration, and scene structure. However, existing attention-level camera encodings either provide ray-only camera signals or rely on pinhole camera geometry, limiting their applicability to general camera control under the Unified Camera Model, including wide-angle and fisheye lenses. To address this limitation, we propose Curved Ray Expectation Positional Encoding (CRePE). CRePE represents each image token as a depth-aware positional distribution along its source ray, providing a Unified Camera Model-compatible positional encoding that captures the projected-path geometry induced by wide-angle and fisheye cameras. CRePE is implemented through a Geometric Attention Adapter added to frozen video DiTs, injecting token-wise scene-distance information into selected attention layers and stabilizing it with pseudo supervision from a monocular geometry foundation model. This design leads to more stable camera control and improves several geometry-aware and perceptual-quality metrics, while remaining competitive on video-quality metrics. Controlled positional-encoding ablations show a better overall average rank than a RayRoPE-style endpoint PE baseline, demonstrating the effectiveness of UCM-aware projected-path integration across diverse camera models. Furthermore, by extending the same positional-encoding pathway to external geometry control through Radial MixForcing, CRePE supports external radial-map control for scene-geometry-conditioned generation and source-video motion transfer beyond camera control.

视频生成位置编码相机控制统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。