arXiv:2511.17185cs.CV2025-11被引 4

让视频视角控制更准,细节更清晰。

PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention

  • 用共享查询的跨注意力机制对齐姿态与图像特征
  • 轨迹精度比现有方法高20%,动态场景更稳定
  • 适合需要精准视角编辑的视频生成任务

我们提出 PostCam,一种简洁高效的新型视角视频生成框架,可在动态场景中同时实现优异的细节保留和精确的相机轨迹编辑。现有方法常在基于姿态的控制(缺乏视觉细节)与基于渲染的引导(依赖几何精度)之间权衡。尽管近期有混合尝试,但精确运动与视觉一致性仍难兼顾,主要因跨模态对齐不足。我们认为,鲁棒控制源于多模态信号的深层对齐,而非输入复杂度提升。核心贡献是查询共享的跨注意力机制,将6-DoF姿态与渲染特征投影至统一潜在空间,使模型在去噪过程中自发实现运动线索与像素级引导的内在一致。实验表明,PostCam 在保持高保真视觉细节的同时,轨迹精度较最先进方法提升20%,在复杂动态场景中表现更优。项目主页已公开:https://cccqaq.github.io/PostCam.github.io/

原文摘要 · Abstract (English)

We propose PostCam, a streamlined framework for novel-view video generation that achieves superior detail preservation and precise camera trajectory editing in dynamic scenes. Current methods often struggle with a trade-off between pose-based control, which lacks visual detail, and rendering-based guidance, which is overly sensitive to geometric accuracy. Despite recent hybrid attempts, achieving precise motion and visual consistency remains challenging due to the lack of effective cross-modal alignment. We argue that robust control stems from the deep alignment of multimodal signals rather than increased input complexity. Our core contribution is the Query-Shared Cross-Attention mechanism, which projects 6-DoF poses and rendered features into a unified latent space. This allows the model to spontaneously achieve intrinsic consistency between motion cues and pixel-level guidance during denoising. Experiments demonstrate that PostCam maintains high-fidelity visual details while outperforming state-of-the-art methods by 20% in trajectory precision, exhibiting superior robustness in complex dynamic scenes. Our project webpage is publicly available at: https://cccqaq.github.io/PostCam.github.io/

视频生成视角控制跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。