arXiv:2512.01383cs.CV2025-12中稿 · WACV2026被引 2

轻量级4D点云视频模型,兼顾实时与离线感知,提升机器人环境理解能力。

PointNet4D: A Lightweight 4D Point Cloud Video Backbone for Online and Offline Perception in Robotic Applications

  • 融合Mamba与Transformer的时序融合模块,高效处理变长序列。
  • 在7个数据集9项任务中性能持续领先,尤其在动态场景建模上表现突出。
  • 适用于机器人实时控制与学习系统,如4D扩散策略和模仿学习。

理解随时间演化的动态4D环境(三维空间+时间)对机器人和交互系统至关重要。这些应用需在资源受限条件下实时处理流式点云视频,同时利用过去和当前观测信息。然而,现有4D主干网络多依赖计算量大的时空卷积或Transformer,难以满足实时需求。本文提出PointNet4D,一种专为在线与离线场景优化的轻量级4D主干网络。核心采用混合Mamba-Transformer时序融合块,结合Mamba的高效状态空间建模与Transformer的双向建模能力,可灵活处理不同长度的在线序列。为增强时序理解,引入4DMAP——一种帧级掩码自回归预训练策略,有效捕捉跨帧运动线索。在7个数据集上的9项任务中广泛验证,性能持续提升。进一步构建了两个机器人应用系统:4D Diffusion Policy与4D Imitation Learning,在RoboTwin与HandoverSim基准测试中取得显著进展。

原文摘要 · Abstract (English)

Understanding dynamic 4D environments-3D space evolving over time-is critical for robotic and interactive systems. These applications demand systems that can process streaming point cloud video in real-time, often under resource constraints, while also benefiting from past and present observations when available. However, current 4D backbone networks rely heavily on spatiotemporal convolutions and Transformers, which are often computationally intensive and poorly suited to real-time applications. We propose PointNet4D, a lightweight 4D backbone optimized for both online and offline settings. At its core is a Hybrid Mamba-Transformer temporal fusion block, which integrates the efficient state-space modeling of Mamba and the bidirectional modeling power of Transformers. This enables PointNet4D to handle variable-length online sequences efficiently across different deployment scenarios. To enhance temporal understanding, we introduce 4DMAP, a frame-wise masked auto-regressive pretraining strategy that captures motion cues across frames. Our extensive evaluations across 9 tasks on 7 datasets, demonstrating consistent improvements across diverse domains. We further demonstrate PointNet4D's utility by building two robotic application systems: 4D Diffusion Policy and 4D Imitation Learning, achieving substantial gains on the RoboTwin and HandoverSim benchmarks.

4D点云机器人感知轻量化模型时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。