arXiv:2604.27702cs.CV2026-04

用结构感知采样与注意力机制提升神经辐射场视频压缩成像质量

RayFormer: Modeling Inter- and Intra-Ray Similarity for NeRF-Based Video Snapshot Compressive Imaging

论文配图:RayFormer: Modeling Inter- and Intra-Ray Similarity for NeRF-Based Video Snapshot Compressive Imaging
图 1 · 摘自论文原文
  • 按图像块采样光线,捕捉内容结构特征
  • 引入跨射线与射线内注意力,建模多尺度相似性
  • 适合需要高保真动态场景重建的研究者

视频快照压缩成像(SCI)可从单次快照测量中重建动态场景。近期基于神经辐射场(NeRF)的方法表现出良好重建性能,但通常采用随机光线采样策略,难以捕捉内容结构相似性,导致重建质量受限。为此,我们首先提出一种分块级光线采样策略,以建模内容结构;随后设计跨射线与射线内注意力网络(RayFormer),同时捕捉同一深度空间邻近点间的跨射线相似性,以及沿视角射线相邻点的射线内相关性;最后,得益于分块采样策略,将总变差先验融入目标函数,增强空间平滑性并抑制伪影。在模拟与真实场景中的实验均表明,所提方法达到当前最优(SOTA)重建性能。

原文摘要 · Abstract (English)

Video snapshot compressive imaging (SCI) enables the reconstruction of dynamic scenes from a single snapshot measurement. Recently, NeRF-based methods have shown promising reconstruction performance. However, such methods typically adopt random ray sampling strategies and fail to capture content structural similarities, resulting in limited reconstruction quality. To address these issues, we first propose a patch-level ray sampling strategy to enable the modeling of content structure. Then, we propose an Inter- and Intra-Ray Transformer (RayFormer) to capture the structural similarities, modeling both inter-ray similarities among spatially neighboring points at the same depth and intra-ray correlations between adjacent points along the viewing ray. Finally, benefiting from the patch-level sampling strategy, the total variation prior is incorporated into the objective function to enhance spatial smoothness and suppress artifacts. Experiments in both simulated and real-world scenes demonstrate that the proposed method achieves state-of-the-art (SOTA) reconstruction performance.

NeRF压缩成像注意力机制视频重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。