arXiv:2606.29071physics.med-phcs.SD2026-06

提出一种无需复杂声道耦合的高效声带振动模型,精准模拟持续发声与闭合过程。

An Optimal Contact-Mechanically Consistent and Flow-Separation Adapted Modeling of Vocal Fold Dynamics

论文配图:An Optimal Contact-Mechanically Consistent and Flow-Separation Adapted Modeling of Vocal Fold Dynamics
图 1 · 摘自论文原文
  • 引入额外阻力与结构力,解决阻尼系统中振荡维持难题
  • 基于高速内窥视频数据优化参数,仿真与实测声门面积波形误差低于3%
  • 兼顾生物力学与气动机制,适合声学建模与语音康复研究

单质量-弹簧-阻尼声带模型虽能有效模拟声带振动且结构简单,但若无声腔耦合则难以在存在结构阻尼时维持振荡;现有集中参数模型亦难准确再现发声时的声门闭合。本研究旨在构建一种可靠的简化单自由度发声模型,可在不依赖声腔模型的前提下实现阻尼系统中的持续振荡,并保持符合发声物理规律的声门闭合。研究利用4名正常发声者发长音/i/的高速视频内窥镜数据,通过基于深度学习的图像分割提取声门面积波形(GAWs),并采用粒子群优化法确定模型参数。引入额外阻力以补偿流动分离,产生维持振荡所需的力不平衡;同时在闭合阶段添加外部结构力以维持闭合相。采用四阶龙格-库塔法求解控制方程,显著提升数值稳定性与精度。模型参数针对个体优化后,仿真与实验声门面积波形的归一化误差均低于3%。所提模型准确再现了个体特异性的声带振动及闭合行为,与实验数据高度一致。整体上,该模型提供了一种计算高效的持续发声仿真框架,无需复杂的源-道耦合,却能捕捉发声的关键生物力学与气动机制。

原文摘要 · Abstract (English)

Single mass-spring-damper models of vocal folds have been effective in simulating vocal fold vibrations without added complexity. However, single-degree-of-freedom models cannot sustain oscillation in the presence of structural damping unless source-tract interaction is considered. Moreover, existing lumped models struggle to accurately simulate vocal fold closure during phonation. This study aims to develop a reliable and simplified single-degree-of-freedom model of phonation that can simulate sustained oscillation in a damped system without incorporating a vocal tract model. Additionally, the proposed model maintains vocal fold closure in a manner consistent with the physics of phonation, addressing a longstanding challenge in existing lumped models. High-speed videoendoscopy (HSV) data from four normophonic subjects producing sustained vowel /i/ were used to extract glottal area waveforms (GAWs) via deep learning-based image segmentation for particle swarm optimization of the model parameters. An additional resistance force was incorporated to compensate for flow separation and generate the force imbalance required for sustained oscillation. An external structural force was also added during closure to sustain the closed phase. The 4th-order Runge-Kutta method was used to solve the governing equations with enhanced numerical stability and accuracy. The model parameters were optimized for individual subjects, resulting in normalized errors below 3% between experimental and simulated GAWs. The proposed model accurately reproduced subject-specific vocal fold vibrations and vocal fold closure in agreement with experimental data. Overall, the proposed model provides a computationally efficient framework for simulating sustained phonation without requiring complex source-tract coupling while capturing the key biomechanical and aerodynamic mechanisms of phonation.

声带建模生物力学语音仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。