arXiv:2503.05600cs.CV2025-03被引 6

单模型实现任意分辨率视频渲染与任意比例渐进编码

Progressively Deformable 2D Gaussian Splatting for Video Representation at Arbitrary Resolutions

  • 用可变形2D高斯点表示视频,通过神经微分方程建模时间演化
  • 训练后按D-最优准则剪枝,支持任意比特率渐进传输
  • 每秒渲染超250帧,兼容多尺度连续码率-画质调节,适合视频压缩应用

隐式神经表示(INRs)虽能实现快速视频压缩和高效处理,但单一模型难以跨码率与分辨率实现可扩展解码。现有方法通常依赖重训练或双分支设计,结构化剪枝也无法提供排列无关的渐进传输顺序。受高斯点积显式结构与高效性启发,我们提出D2GV-AR,一种可变形2D高斯视频表示,可在单个模型中实现任意尺度渲染与任意比例渐进编码。将视频划分为固定长度的图像组,每组用一组基准2D高斯基元表示,其时序演化由神经常微分方程建模。训练与渲染时,依据奈奎斯特采样定理进行尺度感知分组,构建跨分辨率的嵌套层次。模型训练完成后,通过D-最优子集目标对基元进行剪枝,实现任意比特率渐进编码。大量实验表明,D2GV-AR渲染速度超过250 FPS,性能达到或超越近期INR基线,支持多尺度连续码率-质量自适应。

原文摘要 · Abstract (English)

Implicit neural representations (INRs) enable fast video compression and effective video processing, but a single model rarely offers scalable decoding across rates and resolutions. In practice, multi-resolution typically relies on retraining or multi-branch designs, and structured pruning failed to provide a permutation-invariant progressive transmission order. Motivated by the explicit structure and efficiency of Gaussian splatting, we propose D2GV-AR, a deformable 2D Gaussian video representation that enables \emph{arbitrary-scale} rendering and \emph{any-ratio} progressive coding within a single model. We partition each video into fixed-length Groups of Pictures and represent each group with a canonical set of 2D Gaussian primitives, whose temporal evolution is modeled by a neural ordinary differential equation. During training and rendering, we apply scale-aware grouping according to Nyquist sampling theorem to form a nested hierarchy across resolutions. Once trained, primitives can be pruned via a D-optimal subset objective to enable any-ratio progressive coding. Extensive experiments show that D2GV-AR renders at over 250 FPS while matching or surpassing recent INR baselines, enabling multiscale continuous rate--quality adaptation.

视频生成高斯点积渐进编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。