arXiv:2604.11637cs.CV2026-04中稿 · CVPR被引 1

通过频谱分析提升4D点云视频的几何与动态理解能力

STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding

  • 将4D点云视频转为图谱信号,分频段捕捉不同几何结构
  • 在动作识别与语义分割任务上超越现有方法,显著提升性能
  • 适合从事点云视频理解、三维动态建模的研究者

4D点云视频能够捕捉场景丰富的空间与时间动态,在多种4D理解任务中具有独特价值。然而,现有方法大多在时空域进行处理,难以有效捕捉4D点云视频的底层几何特征,导致表征学习与理解能力下降。本文从互补的频谱视角出发,将4D点云视频转换为图谱信号,分解为多个频率带,分别捕捉不同的几何结构。谱分析表明,低频信号保留更粗粒度形状,高频信号编码更精细几何细节。基于此,我们提出统一框架STS-Mixer,融合空间、时间与频谱表示,实现对4D点云视频丰富几何与时间动态的联合建模,支持细粒度与整体化理解。大量实验表明,STS-Mixer在3D动作识别与4D语义分割多个主流基准上均取得持续领先性能。代码与模型已开源。

原文摘要 · Abstract (English)

4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most existing methods work in the spatiotemporal domain where the underlying geometric characteristics of 4D point cloud videos are hard to capture, leading to degraded representation learning and understanding of 4D point cloud videos. We address the above challenge from a complementary spectral perspective. By transforming 4D point cloud videos into graph spectral signals, we can decompose them into multiple frequency bands each of which captures distinct geometric structures of point cloud videos. Our spectral analysis reveals that the decomposed low-frequency signals capture more coarse shapes while high-frequency signals encode more fine-grained geometry details. Building on these observations, we design Spatio-Temporal-Spectral Mixer (STS-Mixer), a unified framework that mixes spatial, temporal, and spectral representations of point cloud videos. STS-Mixer integrates multi-band delineated spectral signals with spatiotemporal information to capture rich geometries and temporal dynamics, while enabling fine-grained and holistic understanding of 4D point cloud videos. Extensive experiments show that STS-Mixer achieves superior performance consistently across multiple widely adopted benchmarks on both 3D action recognition and 4D semantic segmentation tasks. Code and models are available at https://github.com/Vegetebird/STS-Mixer.

4D点云频谱分析动作识别语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。