arXiv:2602.15181cs.CVcs.LG2026-02被引 1

让体育和演出视频可回放任意时刻的视角,实现动态场景的时空一致渲染。

Time-Archival Camera Virtualization for Sports and Visual Performances

  • 用神经体积渲染建模多摄像机同步下的刚性运动,避免点云依赖
  • 支持任意时间点回放并生成新视角,实现动态事件的时序存档
  • 适合体育直播、舞台表演等需快速重播与分析的场景

相机虚拟化作为新型视图合成技术,在视觉娱乐、现场演出和体育转播中具有变革潜力,可通过有限个校准的静态物理摄像机图像生成新视角的逼真图像。尽管近期进展显著,但在快速运动的体育赛事和舞台表演中,实现空间与时间上一致且逼真的动态场景渲染,并具备高效的时间存档能力仍具挑战。现有基于3D高斯溅射(3DGS)的方法虽可实现实时视图合成,但依赖结构光束法获取精确3D点云,难以处理大范围非刚性、快速运动(如翻腾、跳跃、关节动作、球员突变切换),且多个主体独立运动会破坏4DGS、ST-GS等动态溅射方法常用的高斯追踪假设。本文提出重新思考神经体积渲染在相机虚拟化中的应用,以支持高效的时间存档功能,使用户可回溯任意历史时刻的动态场景并进行新视角合成,实现直播事件的回放、分析与存档,这一功能在现有神经渲染与视图合成方法中尚属空白。

原文摘要 · Abstract (English)

Camera virtualization -- an emerging solution to novel view synthesis -- holds transformative potential for visual entertainment, live performances, and sports broadcasting by enabling the generation of photorealistic images from novel viewpoints using images from a limited set of calibrated multiple static physical cameras. Despite recent advances, achieving spatially and temporally coherent and photorealistic rendering of dynamic scenes with efficient time-archival capabilities, particularly in fast-paced sports and stage performances, remains challenging for existing approaches. Recent methods based on 3D Gaussian Splatting (3DGS) for dynamic scenes could offer real-time view-synthesis results. Yet, they are hindered by their dependence on accurate 3D point clouds from the structure-from-motion method and their inability to handle large, non-rigid, rapid motions of different subjects (e.g., flips, jumps, articulations, sudden player-to-player transitions). Moreover, independent motions of multiple subjects can break the Gaussian-tracking assumptions commonly used in 4DGS, ST-GS, and other dynamic splatting variants. This paper advocates reconsidering a neural volume rendering formulation for camera virtualization and efficient time-archival capabilities, making it useful for sports broadcasting and related applications. By modeling a dynamic scene as rigid transformations across multiple synchronized camera views at a given time, our method performs neural representation learning, providing enhanced visual rendering quality at test time. A key contribution of our approach is its support for time-archival, i.e., users can revisit any past temporal instance of a dynamic scene and can perform novel view synthesis, enabling retrospective rendering for replay, analysis, and archival of live events, a functionality absent in existing neural rendering approaches and novel view synthesis...

相机虚拟化动态渲染体育转播时间存档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。