arXiv:2603.16154cs.CVcs.AI2026-03

提出GATS框架,让4D点云视频理解更抗帧率变化和分布不均。

GATS: Gaussian Aware Temporal Scaling Transformer for Invariant 4D Spatio-Temporal Point Cloud Representation

  • 用高斯统计与不确定性门控增强点云邻域聚合
  • 引入可学习时间缩放因子,实现不同帧率下运动估计一致
  • 适合做动态环境感知的点云视频任务,尤其在噪声和遮挡下表现好

理解4D点云视频对智能体感知动态环境至关重要。然而,不同帧率带来的时序尺度偏差以及不规则点云的分布不确定性,使得设计统一且鲁棒的4D主干网络极具挑战。现有基于CNN或Transformer的方法受限于有限感受野或二次计算复杂度,且忽略这些隐式失真。为此,我们提出新型双不变框架——高斯感知时序缩放变换器(GATS),显式解决分布不一致与时序问题。提出的不确定性引导高斯卷积(UGGC)融合局部高斯统计与不确定性感知门控,实现密度变化、噪声和遮挡下的鲁棒邻域聚合。同时,时序缩放注意力(TSA)引入可学习缩放因子以归一化时序距离,确保帧划分不变性及跨帧率的一致速度估计。二者互补:时序缩放先归一化时间间隔,再进行高斯建模,提升对不规则分布的鲁棒性。在主流基准上,MSR-Action3D(+6.62%准确率)、NTU RGBD(+1.4%准确率)和Synthia4D(+1.8% mIoU)实验表明,GATS显著提升性能,相比基于Transformer的方法更具准确性、鲁棒性和可扩展性。

原文摘要 · Abstract (English)

Understanding 4D point cloud videos is essential for enabling intelligent agents to perceive dynamic environments. However, temporal scale bias across varying frame rates and distributional uncertainty in irregular point clouds make it highly challenging to design a unified and robust 4D backbone. Existing CNN or Transformer based methods are constrained either by limited receptive fields or by quadratic computational complexity, while neglecting these implicit distortions. To address this problem, we propose a novel dual invariant framework, termed \textbf{Gaussian Aware Temporal Scaling (GATS)}, which explicitly resolves both distributional inconsistencies and temporal. The proposed \emph{Uncertainty Guided Gaussian Convolution (UGGC)} incorporates local Gaussian statistics and uncertainty aware gating into point convolution, thereby achieving robust neighborhood aggregation under density variation, noise, and occlusion. In parallel, the \emph{Temporal Scaling Attention (TSA)} introduces a learnable scaling factor to normalize temporal distances, ensuring frame partition invariance and consistent velocity estimation across different frame rates. These two modules are complementary: temporal scaling normalizes time intervals prior to Gaussian estimation, while Gaussian modeling enhances robustness to irregular distributions. Our experiments on mainstream benchmarks MSR-Action3D (\textbf{+6.62\%} accuracy), NTU RGBD (\textbf{+1.4\%} accuracy), and Synthia4D (\textbf{+1.8\%} mIoU) demonstrate significant performance gains, offering a more efficient and principled paradigm for invariant 4D point cloud video understanding with superior accuracy, robustness, and scalability compared to Transformer based counterparts.

4D点云时空建模时序不变性点云视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。