用软序数回归自动评估声带损伤严重程度,媲美专家判断。
Classifying Phonotrauma Severity from Vocal Fold Images with Soft Ordinal Regression
- 采用软序数回归处理声带损伤的有序标签,融合标注者不确定性。
- 性能接近临床专家,且能输出可信的置信度估计。
- 适合语音医学研究与大规模临床数据分析场景。
声带创伤是由发声时的力学作用导致的声带组织损伤,其严重程度呈连续谱分布,治疗方案随严重程度变化。当前评估依赖医生主观判断,成本高且一致性差。本文首次提出从声带图像自动分类声带创伤严重程度的方法。为处理标签的有序特性,采用经典的序数回归框架;为应对标注不确定性,提出一种新型损失函数,可处理反映标注者评分分布的软标签。所提软序数回归方法预测性能接近临床专家水平,同时提供校准良好的不确定性估计。该自动化工具有助于开展大规模声带创伤研究,推动临床认知提升与患者诊疗优化。
原文摘要 · Abstract (English)
Phonotrauma refers to vocal fold tissue damage resulting from exposure to forces during voicing. It occurs on a continuum from mild to severe, and treatment options can vary based on severity. Assessment of severity involves a clinician's expert judgment, which is costly and can vary widely in reliability. In this work, we present the first method for automatically classifying phonotrauma severity from vocal fold images. To account for the ordinal nature of the labels, we adopt a widely used ordinal regression framework. To account for label uncertainty, we propose a novel modification to ordinal regression loss functions that enables them to operate on soft labels reflecting annotator rating distributions. Our proposed soft ordinal regression method achieves predictive performance approaching that of clinical experts, while producing well-calibrated uncertainty estimates. By providing an automated tool for phonotrauma severity assessment, our work can enable large-scale studies of phonotrauma, ultimately leading to improved clinical understanding and patient care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。