arXiv:2412.15054cs.CVcs.AI2024-12被引 2

首个公开的声带裂隙语义分割数据集,助力语音康复技术发展

GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation

  • 构建65段喉部高速视频,标注声带裂隙区域
  • 包含健康者、患者及未知状态共50名受试者数据
  • 支持自动分割算法开发,适合语音医学研究者

当前基于高速喉镜视频的辅助播放技术发展受限于缺乏公开标注的语义分割数据集,尤其缺少声门裂区域的精确标注。为解决这一问题,本文提出GIRAFE数据仓库,包含来自50名受试者(30名女性,20名男性)的65段高速喉镜视频,涵盖15例健康对照、26例已确诊嗓音障碍患者以及24例健康状况未知者。所有视频均由专家手动标注声门裂的语义分割掩码,并整合了多种前沿自动分割方法的结果。该数据集已支持多项研究,验证其在开发新型声门裂分割算法中的实用性。尽管已有进展,实现完全自动、精准的声门区域语义分割仍是开放挑战。

原文摘要 · Abstract (English)

The advances in the development of Facilitative Playbacks extracted from High-Speed videoendoscopic sequences of the vocal folds are hindered by a notable lack of publicly available datasets annotated with the semantic segmentations corresponding to the area of the glottal gap. This fact also limits the reproducibility and further exploration of existing research in this field. To address this gap, GIRAFE is a data repository designed to facilitate the development of advanced techniques for the semantic segmentation, analysis, and fast evaluation of High-Speed videoendoscopic sequences of the vocal folds. The repository includes 65 high-speed videoendoscopic recordings from a cohort of 50 patients (30 female, 20 male). The dataset comprises 15 recordings from healthy controls, 26 from patients with diagnosed voice disorders, and 24 with an unknown health condition. All of them were manually annotated by an expert, including the masks corresponding to the semantic segmentation of the glottal gap. The repository is also complemented with the automatic segmentation of the glottal area using different state-of-the-art approaches. This data set has already supported several studies, which demonstrates its usefulness for the development of new glottal gap segmentation algorithms from High-Speed-Videoendoscopic sequences to improve or create new Facilitative Playbacks. Despite these advances and others in the field, the broader challenge of performing an accurate and completely automatic semantic segmentation method of the glottal area remains open.

声带分析数据集医学影像分割算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。