arXiv:2409.03597cs.SDcs.AI2024-09中稿 · Computer Speech an…被引 2

融合音视频分析,自动识别声带麻痹关键片段与特征。

Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Fold Paralysis

  • 结合音视频数据,用扩散模型优化声门分割以减少误检。
  • 在真实临床数据上实现单侧声带麻痹的准确分类。
  • 为医生提供客观量化指标和可视化辅助诊断支持。

本文提出多模态喉镜视频分析系统(MLVAS),利用音频与视频数据自动提取原始喉镜闪光视频中的关键片段与量化指标,辅助临床评估。系统融合基于视频的声门检测与音频关键词定位方法,识别患者发声并精炼视频重点,确保声带运动的最佳观察。除关键片段提取外,MLVAS可生成用于声带麻痹(VFP)检测的有效音视频特征:采用预训练音频编码器提取语音特征;视觉特征通过测量左右声带相对于估计声门中线的夹角偏差获得。为提升分割质量,引入基于扩散模型的后处理模块,对传统U-Net分割结果进行优化,降低假阳性。通过一系列消融实验验证各模块与模态的有效性。在公开分割数据集上的实验表明分割模块有效;在真实临床数据集上的单侧VFP分类结果证明,MLVAS能提供可靠、客观的量化指标与可视化支持,助力临床辅助诊断。

原文摘要 · Abstract (English)

This paper presents the Multimodal Laryngoscopic Video Analyzing System (MLVAS), a novel system that leverages both audio and video data to automatically extract key video segments and metrics from raw laryngeal videostroboscopic videos for assisted clinical assessment. The system integrates video-based glottis detection with an audio keyword spotting method to analyze both video and audio data, identifying patient vocalizations and refining video highlights to ensure optimal inspection of vocal fold movements. Beyond key video segment extraction from the raw laryngeal videos, MLVAS is able to generate effective audio and visual features for Vocal Fold Paralysis (VFP) detection. Pre-trained audio encoders are utilized to encode the patient voice to get the audio features. Visual features are generated by measuring the angle deviation of both the left and right vocal folds to the estimated glottal midline on the segmented glottis masks. To get better masks, we introduce a diffusion-based refinement that follows traditional U-Net segmentation to reduce false positives. We conducted several ablation studies to demonstrate the effectiveness of each module and modalities in the proposed MLVAS. The experimental results on a public segmentation dataset show the effectiveness of our proposed segmentation module. In addition, unilateral VFP classification results on a real-world clinic dataset demonstrate MLVAS's ability of providing reliable and objective metrics as well as visualization for assisted clinical diagnosis.

医学影像声带麻痹多模态分析扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。