arXiv:2604.16836cs.CVcs.AI2026-04

用洛伦兹空间提升语义分割的稳定性与不确定性估计能力

Lorentz Framework for Semantic Segmentation

论文配图:Lorentz Framework for Semantic Segmentation
图 1 · 摘自论文原文
  • 在洛伦兹空间构建无须黎曼优化的分割框架,融合文本与视觉线索
  • 在多个数据集上实现更优分割效果,同时生成置信度图和边界细化
  • 适合需要鲁棒性与不确定性感知的视觉任务,如医疗图像分析

超球空间中的语义分割能紧凑建模层次结构并提供固有的不确定性量化。以往方法多依赖庞加莱球模型,存在数值不稳定性、优化与计算挑战。本文提出一种新型、可操作的、架构无关的语义分割框架(像素级与掩码分类),基于超球洛伦兹模型。利用包含语义与视觉线索的文本嵌入,引导洛伦兹空间中的层次化像素表示。该方法实现稳定高效优化,无需黎曼优化器,并可无缝集成现有欧氏架构。除分割外,本方法还免费提供不确定性估计、置信度图、边界细化、层次化与基于文本的检索及零样本性能,达到更平坦的泛化极小值。我们引入新的洛伦兹锥嵌入不确定性指标。进一步通过梯度分析提供洛伦兹优化的理论与实证洞察。在ADE20K、COCO-Stuff-164k、Pascal-VOC与Cityscapes上,使用DeepLabV3、SegFormer、mask2former与maskformer等先进模型进行大量实验,验证了方法的有效性与通用性。结果表明,超球洛伦兹嵌入在鲁棒且具备不确定性感知的语义分割中具有巨大潜力。代码已开源:https://github.com/mxahan/Lorentz_semantic_segmentation。

原文摘要 · Abstract (English)

Semantic segmentation in hyperbolic space enables compact modeling of hierarchical structure while providing inherent uncertainty quantification. Prior approaches predominantly rely on the Poincaré ball model, which suffers from numerical instability, optimization, and computational challenges. We propose a novel, tractable, architecture-agnostic semantic segmentation framework (pixel-wise and mask classification) in the hyperbolic Lorentz model. We employ text embeddings with semantic and visual cues to guide hierarchical pixel-level representations in Lorentz space. This enables stable and efficient optimization without requiring a Riemannian optimizer, and easily integrates with existing Euclidean architectures. Beyond segmentation, our approach yields free uncertainty estimation, confidence map, boundary delineation, hierarchical and text-based retrieval, and zero-shot performance, reaching generalized flatter minima. We introduce a novel uncertainty and confidence indicator in Lorentz cone embeddings. Further, we provide analytical and empirical insights into Lorentz optimization via gradient analysis. Extensive experiments on ADE20K, COCO-Stuff-164k, Pascal-VOC, and Cityscapes, utilizing state-of-the-art per-pixel classification models (DeepLabV3 and SegFormer) and mask classification models (mask2former and maskformer), validate the effectiveness and generality of our approach. Our results demonstrate the potential of hyperbolic Lorentz embeddings for robust and uncertainty-aware semantic segmentation. Code is available at https://github.com/mxahan/Lorentz_semantic_segmentation.

语义分割超球几何不确定性估计洛伦兹空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。