用分层动作引导的动态世界模型,提升耳部结构精细分割准确率
AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans

- 基于递归状态空间模型,通过多步潜变量演进实现分层解剖推理
- 对小而不规则、重叠结构的分割性能提升超43%,HD95显著降低
- 适合处理复杂解剖结构与嵌套标签的医学图像分割任务
CT中耳部结构的精细分割面临挑战:耳部占图像区域小,软骨边界高度不规则,软骨与周围软组织界面模糊。临床标注包含包含软骨和邻近皮肤的复合结构及其对应的纯软骨区域,导致嵌套与重叠标签。我们提出一种基于世界模型的分割框架,实现超越传统前馈预测的迭代解剖推理。该框架基于编码器-解码器结构,在中间潜空间引入确定性递归状态空间模型。多尺度编码特征与部分解码表示融合形成结构观测,初始化潜动力学。推理时,模型在无真值指导下执行三步潜变量滚动。分层解剖动作更新递归状态,逐步精炼潜表示。最终潜轨迹投影回解码器,结合高分辨率特征生成分割结果。为学习可靠潜转移,引入平衡的分层动作目标,解决前景稀疏、解剖组缺失及增删操作失衡问题。大量实验表明,该框架在小、不规则、重叠耳部结构上的分割精度持续提升,HD95降低超过43%。结果验证了潜世界模型推理在挑战性医学图像分割中的有效性。
原文摘要 · Abstract (English)
Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly irregular, and interfaces between cartilage and surrounding soft tissues are often ambiguous. Clinical annotations may also include both composite structures containing cartilage and adjacent skin and their corresponding cartilage-only regions, producing nested and overlapping labels. We propose a world-model-based segmentation framework that enables iterative anatomical reasoning beyond conventional feed-forward prediction. Built on an encoder-decoder architecture, the framework introduces a deterministic recurrent state-space model into the intermediate latent space. Multi-scale encoder features and partially decoded representations are fused to form a structural observation that initializes the latent dynamics. During inference, the model performs a three-step latent rollout without ground-truth guidance. Hierarchical anatomical actions update the recurrent state and progressively refine the latent representation. The resulting latent trajectory is projected back into the decoder and combined with high-resolution features to produce the final segmentation. To learn reliable latent transitions, we introduce a balanced hierarchical action objective that addresses foreground sparsity, missing anatomical groups, and imbalance between add and remove operations. Extensive experiments show that the proposed framework consistently improves segmentation accuracy and reduces HD95 by more than 43% for small, irregular, and overlapping auricular structures in CT. These results demonstrate the effectiveness of latent world-model reasoning for challenging medical image segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。