arXiv:2607.28935cs.CV2026-07

解决室内语义占据的长尾分布问题,提升小类别预测精度

Group-wise Supervision with Focal-Dice Loss for Long-Tailed Indoor Semantic Occupancy Prediction

论文配图:Group-wise Supervision with Focal-Dice Loss for Long-Tailed Indoor Semantic Occupancy Prediction
图 1 · 摘自论文原文
  • 按语义分组设计多尺度并行预测头,增强模型对尾部类别的学习能力
  • 提出统一焦点-狄利克雷损失,动态聚焦难样本并保持几何完整性
  • 在EmbodiedScan数据集上显著提升长尾类别性能,相对基线增益11.38%

近年来,3D语义占据预测在理解室内场景方面受到越来越多关注。然而,与结构化的室外环境不同,室内场景包含大量类别且呈现严重的长尾分布,已成为制约现有模型性能的核心瓶颈。为此,我们提出一种新方法Group-UFD Occ,基于分层语义监督与协同损失优化。在架构层面,引入细粒度语义分组策略,并设计多尺度、并行的“主专家”预测头,通过深度正则化引导模型高效学习尾部类别特征。在优化层面,提出统一焦点-狄利克雷(UFD)损失函数,在体素级动态聚焦难样本的同时,从区域视角同步优化预测物体的几何完整性。我们在大规模EmbodiedScan数据集上进行了实验,结果表明,该方法相比基线相对提升11.38%,并在多个关键长尾类别上实现显著精度提升。

原文摘要 · Abstract (English)

Recently, 3D semantic occupancy prediction has garnered increasing attention for understanding the indoor scene. However, unlike structured outdoor environments, indoor scenes feature a high diversity of object categories that exhibit a severe long-tailed distribution, which has become a core bottleneck limiting the performance of existing models. To tackle this challenge, we propose a novel method, Group-UFD Occ, based on hierarchical semantic supervision and synergistic loss optimization. At the architectural level, we introduce a fine-grained semantic grouping strategy and design multi-scale, parallel ``main-expert'' prediction heads to guide the model in efficiently learning tail-class features through deep regularization. At the optimization level, we introduce the Unified Focal-Dice (UFD) loss. This synergistic loss function dynamically focuses on hard samples at the per-voxel level. Meanwhile, it simultaneously optimizes the geometric integrity of predicted objects from a region-based perspective. We conducted experiments on the large-scale EmbodiedScan dataset. The results demonstrate that our method yields a relative improvement of 11.38\% over the baseline, with substantial accuracy gains in several critical long-tailed categories.

3D语义长尾分布占据预测损失函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。