arXiv:2509.15473eess.AScs.CL2025-09被引 3

通过呼吸与语义停顿检测,评估运动后身体恢复状态。

Breathing and Semantic Pause Detection and Exertion-Level Classification in Post-Exercise Speech

  • 基于音频与呼吸信号同步数据,标注三类停顿类型。
  • 语义停顿检测准确率达89%,整体分类准确率73%。
  • 适合运动康复、语音生理分析领域研究者参考。

运动后语音富含生理与语言线索,常表现为语义停顿、呼吸停顿及呼吸-语义复合停顿。识别这些事件可评估恢复速度、肺功能及运动负荷异常。然而,现有研究在该场景下对不同类型停顿的识别与区分仍有限。本文基于新发布的同步音频与呼吸信号数据集,提供系统化的停顿类型标注。利用这些标注,系统性地探索了多种深度学习模型(GRU、1D CNN-LSTM、AlexNet、VGG16)、声学特征(MFCC、MFB)以及分层Wav2Vec2表示下的呼吸与语义停顿检测及运动强度分类。评估三种设置——单特征、特征融合、两级检测-分类级联——在分类与回归框架下的表现。结果表明,各类停顿检测准确率最高达89%(语义)、55%(呼吸)、86%(复合停顿),整体准确率为73%;运动强度分类达到90.5%准确率,优于先前工作。

原文摘要 · Abstract (English)

Post-exercise speech contains rich physiological and linguistic cues, often marked by semantic pauses, breathing pauses, and combined breathing-semantic pauses. Detecting these events enables assessment of recovery rate, lung function, and exertion-related abnormalities. However, existing works on identifying and distinguishing different types of pauses in this context are limited. In this work, building on a recently released dataset with synchronized audio and respiration signals, we provide systematic annotations of pause types. Using these annotations, we systematically conduct exploratory breathing and semantic pause detection and exertion-level classification across deep learning models (GRU, 1D CNN-LSTM, AlexNet, VGG16), acoustic features (MFCC, MFB), and layer-stratified Wav2Vec2 representations. We evaluate three setups-single feature, feature fusion, and a two-stage detection-classification cascade-under both classification and regression formulations. Results show per-type detection accuracy up to 89$\%$ for semantic, 55$\%$ for breathing, 86$\%$ for combined pauses, and 73$\%$overall, while exertion-level classification achieves 90.5$\%$ accuracy, outperformin prior work.

语音分析运动生理停顿检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。