用扩散模型融合多焦点信息,提升胚胎发育阶段识别准确率
EmbryoDiff: A Conditional Diffusion Framework with Multi-Focal Feature Fusion for Fine-Grained Embryo Developmental Stage Recognition
- 将胚胎阶段识别建模为条件去噪过程,融合多焦点特征构建3D形态表征
- 单步去噪即达82.8%和81.3%准确率,优于现有方法
- 适合辅助IVF临床决策,尤其在细胞遮挡场景下表现稳健
体外受精(IVF)中精细识别胚胎发育阶段对评估胚胎活力至关重要。尽管深度学习方法已取得良好效果,但现有判别模型未能利用胚胎发育的分布先验,且依赖单一焦点信息导致表征不完整,在细胞遮挡下易产生特征歧义。为此,我们提出EmbryoDiff,一种两阶段扩散框架,将任务建模为条件序列去噪过程。首先训练并冻结帧级编码器提取鲁棒的多焦点特征;第二阶段引入多焦点特征融合策略,跨焦平面聚合信息以构建3D感知的形态表示,有效缓解遮挡引起的歧义。基于该融合表征,提取互补的语义与边界线索,并设计混合语义-边界条件模块注入扩散去噪过程,实现精准分类。在两个基准数据集上的实验表明,本方法达到领先性能。值得注意的是,仅需一步去噪,模型即分别取得82.8%和81.3%的平均测试准确率。
原文摘要 · Abstract (English)
Identification of fine-grained embryo developmental stages during In Vitro Fertilization (IVF) is crucial for assessing embryo viability. Although recent deep learning methods have achieved promising accuracy, existing discriminative models fail to utilize the distributional prior of embryonic development to improve accuracy. Moreover, their reliance on single-focal information leads to incomplete embryonic representations, making them susceptible to feature ambiguity under cell occlusions. To address these limitations, we propose EmbryoDiff, a two-stage diffusion-based framework that formulates the task as a conditional sequence denoising process. Specifically, we first train and freeze a frame-level encoder to extract robust multi-focal features. In the second stage, we introduce a Multi-Focal Feature Fusion Strategy that aggregates information across focal planes to construct a 3D-aware morphological representation, effectively alleviating ambiguities arising from cell occlusions. Building on this fused representation, we derive complementary semantic and boundary cues and design a Hybrid Semantic-Boundary Condition Block to inject them into the diffusion-based denoising process, enabling accurate embryonic stage classification. Extensive experiments on two benchmark datasets show that our method achieves state-of-the-art results. Notably, with only a single denoising step, our model obtains the best average test performance, reaching 82.8% and 81.3% accuracy on the two datasets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。