用SAM 2零样本分割瞳孔,1400万张眼图表现媲美专用模型。
Zero-Shot Pupil Segmentation with SAM 2: A Case Study of Over 14 Million Images
- 仅需点击一次即可实现跨数据集的零样本瞳孔分割。
- 在1400万张图像上达到93% mIoU,无需微调。
- 适合眼动追踪研究者快速部署,降低标注成本。
我们探索了SAM 2这一视觉基础模型在提升注视估计与眼动追踪技术方面的变革潜力。通过大幅减少标注时间、降低部署门槛并提高分割精度,SAM 2有效应对了研究人员和实践者面临的关键挑战。利用其零样本分割能力,仅需每视频点击一次,我们在超过1400万张来自多样化数据集的眼部图像上测试了SAM 2,涵盖虚拟现实场景及全球最大的可穿戴眼动仪统一数据集。令人惊讶的是,在瞳孔分割任务中,SAM 2在未微调的情况下实现了与专用于眼部图像训练的模型相当的性能,达到高达93%的平均交并比(mIoU)。此外,我们公开了代码与这些常用数据集的分割掩码,以推动后续研究。
原文摘要 · Abstract (English)
We explore the transformative potential of SAM 2, a vision foundation model, in advancing gaze estimation and eye tracking technologies. By significantly reducing annotation time, lowering technical barriers through its ease of deployment, and enhancing segmentation accuracy, SAM 2 addresses critical challenges faced by researchers and practitioners. Utilizing its zero-shot segmentation capabilities with minimal user input-a single click per video-we tested SAM 2 on over 14 million eye images from diverse datasets, including virtual reality setups and the world's largest unified dataset recorded using wearable eye trackers. Remarkably, in pupil segmentation tasks, SAM 2 matches the performance of domain-specific models trained solely on eye images, achieving competitive mean Intersection over Union (mIoU) scores of up to 93% without fine-tuning. Additionally, we provide our code and segmentation masks for these widely used datasets to promote further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。