跨传感器森林点云树实例分割新框架,精度高且泛化能力强。
SegmentAnyTreeV2: Scaling Transformer-Based Tree Instance Segmentation Across Sensors, Platforms, and Forests

- 基于点变换器的序列化结构,结合树聚焦注意力解码器。
- 在427个场景上实现90.5%精度、85.0% F1,优于现有方法。
- 适用于多平台、多林区,零样本迁移能力出色。
我们提出SegmentAnyTreeV2,一种传感器与平台无关的森林点云语义与实例分割框架。模型采用基于序列化的Point Transformer v3主干,搭配轻量级语义头和树聚焦的交叉注意力掩码解码器。语义预测将实例解码限制在树类体素内,通过实例感知查询初始化、一对多种子监督和非对称掩码评分,提升密集与复杂结构林分的分离效果。我们还构建了FOR-instance v3,一个包含427个场景和26,496棵标注树木的扩展基准,覆盖多种生物群落、林分结构与LiDAR平台。在FOR-instanceV2测试集上,SegmentAnyTreeV2达到90.5%精度、80.2%召回率、85.0% F1、90.7%覆盖率和87.6%语义mIoU,显著优于以往学习型方法,在实例检测与掩码完整性方面表现更优。在独立站点的零样本评估中也展现出强跨域泛化能力。
原文摘要 · Abstract (English)
We present SegmentAnyTreeV2, a sensor- and platform-agnostic framework for semantic and instance segmentation of forest point clouds. The model combines a serialization-based Point Transformer v3 backbone with a lightweight semantic head and a tree-focused cross-attention mask decoder. Semantic predictions restrict instance decoding to tree-class voxels, while instance-aware query initialization, one-to-many seed supervision, and asymmetric mask scoring improve separation in dense and structurally complex stands. We further introduce FOR-instance v3, an expanded benchmark comprising 427 scenes and 26,496 annotated trees across diverse biomes, forest structures, and LiDAR platforms. On the FOR-instanceV2 test split, SegmentAnyTreeV2 achieves 90.5% precision, 80.2% recall, 85.0% F1, 90.7% coverage, and 87.6% semantic mIoU, outperforming previous learning-based methods in both instance detection and mask completeness. Zero-shot evaluation on independent sites further demonstrates strong cross-domain generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。