arXiv:2607.24591cs.CV2026-07被引 1

实现任意镜头控制的视频重拍,无需3D重建即可自由调节视角和焦距。

CameraAnything: Refilming Videos with Arbitrary Camera Control

论文配图:CameraAnything: Refilming Videos with Arbitrary Camera Control
图 1 · 摘自论文原文
  • 通过像素级射线注入与分辨率感知位置编码,联合控制相机内外参数。
  • 单次生成完成视角切换、焦距调整、分辨率适配等多类编辑操作。
  • 自建合成数据流水线解决真实数据稀缺问题,适合影视制作与跨平台适配。

我们提出 CameraAnything,首个统一的相机可控视频编辑框架,可联合控制相机内参与外参。现有方法或依赖昂贵的3D重建以实现完整相机功能,或仅限于外参修改。由于内参与外参对视频外观的耦合影响,解耦建模尤为困难。为此,我们采用像素级Plücker射线注入,并结合分辨率感知的3D RoPE,在目标潜在表示上构建相机条件与空间位置编码,实现无裁剪、无外推的相机位置、焦距与原始分辨率编辑。为克服成对训练数据稀缺问题,我们进一步设计可扩展的合成数据生成流程,通过结构化多相机录制构建多样动态场景,并生成具有不同相机配置的同步视频。采用定制正交训练策略,CameraAnything 支持在单次生成中实现表达丰富的视频重拍,包括任意视角控制、焦距调节、分辨率自适应及多镜头过渡,为电影级视频编辑与跨平台内容适配提供强大实用价值。

原文摘要 · Abstract (English)

We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters. Existing approaches either rely on expensive 3D reconstruction to achieve full camera functionality or restrict editing to extrinsic parameter manipulation. Moreover, the coupled influence of intrinsic and extrinsic parameters on video appearance makes disentangled modeling particularly challenging. To address this, we adopt per-pixel Plücker ray injection alongside resolution-aware 3D RoPE in self-attention, building both camera conditioning and spatial positional encoding on the target latent to jointly control camera position, focal length, and native resolution editing without cropping or outpainting. To overcome the scarcity of paired training data, we further develop a scalable synthetic pipeline that constructs diverse dynamic scenes through structured multi-camera recording and generates synchronized videos with varied camera configurations. With a tailored orthogonal training strategy, CameraAnything enables expressive video reshooting with arbitrary viewpoint control, focal length adjustment, resolution adaptation, and multi-shot transitions within a single generation process, offering strong practical value for cinematic video editing and cross-platform content adaptation in video production.

视频重拍相机控制生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。