arXiv:2503.14837cs.CVcs.RO2025-03被引 3

联合预测动态场景中物体运动与分割,提升自动驾驶感知精度。

SemanticFlow: A Self-Supervised Framework for Joint Scene Flow Prediction and Instance Segmentation in Dynamic Environments

  • 分阶段联合建模,用粗分割指导精炼运动与语义信息。
  • 在Argoverse和Waymo上实现更高精度的分割与光流估计。
  • 自监督学习无需标注,适合真实复杂交通场景应用。

准确感知动态交通场景对高级自动驾驶系统至关重要,需同时实现鲁棒的物体运动估计与实例分割。传统方法将二者分开处理,导致性能不佳、时空不一致且效率低下。本文提出SemanticFlow多任务框架,联合预测全分辨率点云的场景光流与实例分割。核心创新包括:1)基于粗到精的多任务机制,利用初始粗分割提供上下文信息,通过共享特征模块优化运动与语义;2)设计一组损失函数,提升光流与分割性能,并保证静态与动态物体在时空上的一致性;3)采用自监督学习,利用粗分割识别刚性物体并计算其帧间变换矩阵,生成自监督标签。在Argoverse和Waymo数据集上验证,该框架在实例分割精度、场景光流估计及计算效率方面均优于现有自监督方法,树立了动态场景理解的新基准。

原文摘要 · Abstract (English)

Accurate perception of dynamic traffic scenes is crucial for high-level autonomous driving systems, requiring robust object motion estimation and instance segmentation. However, traditional methods often treat them as separate tasks, leading to suboptimal performance, spatio-temporal inconsistencies, and inefficiency in complex scenarios due to the absence of information sharing. This paper proposes a multi-task SemanticFlow framework to simultaneously predict scene flow and instance segmentation of full-resolution point clouds. The novelty of this work is threefold: 1) developing a coarse-to-fine prediction based multi-task scheme, where an initial coarse segmentation of static backgrounds and dynamic objects is used to provide contextual information for refining motion and semantic information through a shared feature processing module; 2) developing a set of loss functions to enhance the performance of scene flow estimation and instance segmentation, while can help ensure spatial and temporal consistency of both static and dynamic objects within traffic scenes; 3) developing a self-supervised learning scheme, which utilizes coarse segmentation to detect rigid objects and compute their transformation matrices between sequential frames, enabling the generation of self-supervised labels. The proposed framework is validated on the Argoverse and Waymo datasets, demonstrating superior performance in instance segmentation accuracy, scene flow estimation, and computational efficiency, establishing a new benchmark for self-supervised methods in dynamic scene understanding.

场景光流实例分割自监督学习自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。