arXiv:2504.00515cs.LGcs.AI2025-04被引 1

用手机图像自动测眼睑参数,DINOv2模型表现更准更省资源。

Training Frozen Feature Pyramid DINOv2 for Eyelid Measurements with Infinite Encoding and Orthogonal Regularization

  • 冻结预训练DINOv2模型,搭配轻量回归器实现高效移动端部署。
  • 结合焦点损失与正交正则化,提升小样本和不平衡数据下的精度。
  • 在手机图像上实现眼睑距离与提肌功能的精准测量,适合临床应用。

准确测量眼睑参数(如MRD1、MRD2和提肌功能LF)对眼整形诊断至关重要,但现有方法依赖人工且结果不一致。本研究评估了SE-ResNet、EfficientNet及基于视觉变换器的DINOv2模型,利用智能手机拍摄图像实现自动化测量。在冻结与微调两种设置下,使用均方误差(MSE)、平均绝对误差(MAE)和决定系数(R²)进行评估。预训练自监督学习的DINOv2展现出更强的可扩展性与鲁棒性,尤其在冻结条件下更适配移动端部署。轻量级回归器(如MLP与深度集成)在极低计算开销下实现高精度。为缓解类别不平衡并提升泛化能力,引入焦点损失、正交正则化和二值编码策略。结果表明,DINOv2结合上述优化,在所有任务中均实现稳定、准确的预测,是面向真实世界、移动友好的临床应用理想选择。该工作凸显了基础模型在推动人工智能眼科诊疗中的潜力。

原文摘要 · Abstract (English)

Accurate measurement of eyelid parameters such as Margin Reflex Distances (MRD1, MRD2) and Levator Function (LF) is critical in oculoplastic diagnostics but remains limited by manual, inconsistent methods. This study evaluates deep learning models: SE-ResNet, EfficientNet, and the vision transformer-based DINOv2 for automating these measurements using smartphone-acquired images. We assess performance across frozen and fine-tuned settings, using MSE, MAE, and R2 metrics. DINOv2, pretrained through self-supervised learning, demonstrates superior scalability and robustness, especially under frozen conditions ideal for mobile deployment. Lightweight regressors such as MLP and Deep Ensemble offer high precision with minimal computational overhead. To address class imbalance and improve generalization, we integrate focal loss, orthogonal regularization, and binary encoding strategies. Our results show that DINOv2 combined with these enhancements delivers consistent, accurate predictions across all tasks, making it a strong candidate for real-world, mobile-friendly clinical applications. This work highlights the potential of foundation models in advancing AI-powered ophthalmic care.

眼睑测量DINOv2移动端AI自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。