arXiv:2608.11037cs.CVcs.IR2026-08

通过多层级证据聚合,提升罕见病面部表型检索准确率

Multi-Level Evidence Aggregation for Robust Facial Phenotype Retrieval in Rare Genetic Disorder Prioritization

论文配图:Multi-Level Evidence Aggregation for Robust Facial Phenotype Retrieval in Rare Genetic Disorder Prioritization
图 1 · 摘自论文原文
  • 在推理阶段融合多张患者图像与疾病中心特征
  • 顶1准确率最高提升至60.94%,跨数据集效果显著
  • 无需重训练模型,适合医疗辅助诊断场景

AI辅助的面部表型分析可通过从如GestaltMatcher数据库(GMDB)等参考库中检索视觉相似的已确诊病例,支持罕见遗传病的优先排序。现有基于GestaltMatcher的检索框架将测试图像与图库中每张图像逐一对比,但这种点对点方法未能充分利用多重证据——患者可能有多张图像,而疾病也可能由多个确诊患者代表。本文提出一种推理时的多层级证据聚合框架,不需修改底层GestaltMatcher-Arc编码器。该框架整合了同一患者多图像的嵌入级聚合、基于患者权重的疾病中心向量,以及个体-中心混合评分机制,融合测试患者观测、疾病级别图库证据和局部最近邻信息。在包含训练中出现疾病(GMDB-Freq)、未见疾病(GMDB-Rare)及多图像患者子集的GMDB v1.1.4上评估。多层级聚合在所有子集上均提升平均每病种前N检索准确率:在GMDB-Freq上顶1准确率从38.52%升至48.82%;在GMDB-Rare上从19.38%升至23.79%;在多图像子集上,GMDB-Multi-Freq从46.12%升至60.94%,GMDB-Multi-Rare从18.54%升至26.71%。结果表明,推理时聚合可有效提升下一代面部表型检索性能,无需重新训练编码器,推动从单图像匹配转向患者与疾病多层级证据整合。

原文摘要 · Abstract (English)

AI-assisted facial phenotyping supports rare genetic disorder prioritization by retrieving visually similar diagnosed cases from facial image reference databases such as the GestaltMatcher Database (GMDB). Existing GestaltMatcher-based retrieval frameworks compare each test image with individual gallery images in a facial phenotype embedding space. However, this pointwise formulation does not fully exploit available evidence, because patients may have multiple images and disorders may be represented by multiple diagnosed gallery patients. We propose an inference-time multi-level evidence aggregation framework that improves facial phenotype retrieval without modifying the underlying GestaltMatcher-Arc encoder. The framework combines embedding-level patient aggregation of multiple images from the same individual, patient-weighted disorder centroids, and hybrid individual-centroid scoring to integrate test-patient observations, disorder-level gallery evidence, and local nearest-neighbor evidence. We evaluated the approach on GMDB v1.1.4 across disorders represented during training (GMDB-Freq), unseen disorders (GMDB-Rare), and multi-image patient subsets, using a unified gallery containing both GMDB-Freq and GMDB-Rare disorders. Multi-level evidence aggregation improved mean per-disorder top-$N$ retrieval accuracy across all evaluation subsets. Top-1 accuracy increased from 38.52% to 48.82% on GMDB-Freq and from 19.38% to 23.79% on GMDB-Rare. On multi-image subsets, top-1 accuracy increased from 46.12% to 60.94% on GMDB-Multi-Freq and from 18.54% to 26.71% on GMDB-Multi-Rare. These findings show that inference-time aggregation can improve next-generation facial phenotype retrieval without retraining the encoder, supporting a shift from isolated single-image matching toward multi-level aggregation of patient and disorder evidence for rare-disorder prioritization.

面部表型罕见病多模态融合医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。