arXiv:2606.08980cs.CV2026-06被引 3

端到端3D全景分割框架,提升语义与实例一致性

EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation

论文配图:EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation
图 1 · 摘自论文原文
  • 直接从多视角图像输出3D语义和实例特征,无需预处理
  • 在Replica数据集上语义分割mIoU提升13%,每场景仅需1秒
  • 适合机器人操作与3D场景编辑等实时应用

本文提出EPS3D,一种全新的开放词汇3D全景分割端到端前馈框架。不同于依赖额外预处理的现有方法,我们设计了端到端架构,并在多样化的3D场景上采用基于知识蒸馏的训练策略,直接从多视图图像预测3D感知的语义与实例特征,提升了3D一致性并避免误差累积。进一步提出互增强模块,通过实例内语义对齐(Ins2Sem)和语义引导实例优化(Sem2Ins),强化语义与实例间的内在一致性,实现更连贯的3D场景理解。最终,EPS3D在两个基准测试中超越当前最优方法(如在Replica上语义部分mIoU提升13%),且效率极高(每场景仅需1秒),支持机器人操作与3D场景编辑等任务。

原文摘要 · Abstract (English)

This paper introduces EPS3D, a new end-to-end feed-forward framework for open-vocabulary 3D panoptic segmentation. Unlike existing methods relying on additional preprocessing, we design an end-to-end architecture, with a distillation-based training strategy on diverse 3D scenes to predict 3D-aware semantic and instance features from multi-view images, improving 3D consistency and avoiding error accumulation. We further propose a mutual enhancement module to enforce inherent semantic-instance consistency. By aligning semantics within instances (Ins2Sem) and refining instance features with semantic guidance (Sem2Ins), we achieve more coherent 3D scene understanding. Ultimately, EPS3D outperforms SOTA baselines on two benchmarks (e.g., +13% mIoU for semantics on Replica) with high efficiency (e.g., 1s per scene), supporting tasks like robotic manipulation and 3D scene editing.

3D分割全景分割多视图实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。