基于面部图像的医学嵌入模型,提升生物特征与医疗分析精度。
MeFEm: Medical Face Embedding model
- 改进JEPA架构,采用轴向条纹掩码聚焦语义区域。
- 在少数据下超越FaRL、Franca等基线,准确率更高。
- 适合医学影像分析与生物特征识别研究者使用。
我们提出MeFEm,一种基于改进联合嵌入预测架构(JEPA)的视觉模型,用于从面部图像中进行生物特征与医学分析。关键改进包括轴向条纹掩码策略,以聚焦于语义相关区域;圆形损失加权方案;以及对CLS token的概率重分配,以提升线性探针质量。模型在整合后的精选图像数据集上训练,尽管使用数据量显著较少,但在核心人体测量任务上仍优于强基线模型如FaRL和Franca。此外,在新型整合的封闭源数据集上评估的体质量指数(BMI)估计任务中也表现出色,有效缓解了现有数据中的领域偏差问题。模型权重已公开于https://huggingface.co/boretsyury/MeFEm,为该领域未来研究提供有力基准。
原文摘要 · Abstract (English)
We present MeFEm, a vision model based on a modified Joint Embedding Predictive Architecture (JEPA) for biometric and medical analysis from facial images. Key modifications include an axial stripe masking strategy to focus learning on semantically relevant regions, a circular loss weighting scheme, and the probabilistic reassignment of the CLS token for high quality linear probing. Trained on a consolidated dataset of curated images, MeFEm outperforms strong baselines like FaRL and Franca on core anthropometric tasks despite using significantly less data. It also shows promising results on Body Mass Index (BMI) estimation, evaluated on a novel, consolidated closed-source dataset that addresses the domain bias prevalent in existing data. Model weights are available at https://huggingface.co/boretsyury/MeFEm , offering a strong baseline for future work in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。