从频域角度建模点云动作,同时捕捉全局运动与精细时序细节。
SRENet: Spectral Re-Entry Network for Point Cloud Action Recognition

- 通过频域分解块分离高低频特征,实现时空动态解耦
- 在三个数据集上达到当前最佳性能,尤其在复杂动作识别上提升显著
- 适合需要高精度动作理解的自动驾驶与人机交互场景
从点云序列中识别人类动作对自动驾驶和人机交互等三维感知应用至关重要。然而,点云的不规则结构和时间不一致性给时空表示学习带来挑战,尤其在捕捉全局运动上下文与细粒度时序动态方面。本文提出SRENet,一种基于频域感知的框架,从频率视角显式学习运动的全局上下文与细粒度时序动态。SRENet引入频域分解块(SDeBlock),沿时空轴进行小波分析,将特征解耦为低频与高频成分,并施加频域特定注意力。为恢复残差动态并重新对齐语义融合中扭曲的时间频域结构,设计了频域重进入块(SReBlock)进行二次时间分解。此外,提出频域感知学习策略,通过对比损失与渐进式课程调度,逐步聚焦从低频到高频空间,契合粗粒度到细粒度运动模式。在MSR-Action3D、NTU-RGBD和NTU-RGBD120三个数据集上的大量实验表明,SRENet达到当前最优性能,验证了频域建模在点云动作理解中的有效性。
原文摘要 · Abstract (English)
Recognizing human actions from point cloud sequences is critical for 3D perception driven applications such as autonomous driving and human-computer interaction. However, the irregular structure and temporal inconsistency of point clouds pose unique challenges for spatio-temporal representation learning, especially in capturing both global motion context and fine-grained temporal dynamics. We propose SRENet, a spectral-aware framework designed to explicitly learn both global context and fine-grained temporal dynamics of motion from a frequency perspective for action recognition. SRENet introduces a Spectral Decomposition Block (SDeBlock) that performs wavelet-based analysis along temporal and spatial axes, disentangling features into low- and high-frequency components with frequency-specific attention. To recover residual dynamics and re-align temporal frequency structures distorted during semantic fusion, a Spectral Re-entry Block (SReBlock) performs secondary temporal decomposition. Furthermore, a spectral-aware learning strategy is devised to enhance discriminability in both frequency subspaces via contrastive loss and a curriculum schedule that gradually shifts focus from low- to high-frequency spaces in line with coarse to detailed motion patterns. Extensive experiments on MSR-Action3D, NTU-RGBD and NTU-RGBD120 demonstrate that SRENet achieves state-of-the-art performance, validating the effectiveness of frequency modeling in point cloud-based action understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。