arXiv:2607.18629cs.SDeess.SP2026-07

用混沌物理指导肌电转语音,模型更小性能更强。

CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses

论文配图:CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses
图 1 · 摘自论文原文
  • 基于Samba架构与混沌损失函数,融合非线性动力学特性
  • 参数量减少40.79%,语音质量指标提升2.1至4.7倍
  • 首次将混沌物理用于语音合成,适合低资源语音重建场景

我们提出一种受混沌启发的新型肌电转语音(EMG-to-Speech, ETS)架构CS-ETS,结合基于Samba的编码器与两种新颖的混沌启发损失函数——李雅普诺夫指数正则化(LER)和多尺度去趋势波动分析(MSDFA)。LER基于李雅普诺夫指数,捕捉非线性波动与对初值的敏感性;MSDFA利用去趋势波动分析量化分形式、长程时间混沌相关性。CS-ETS在参数量上比先前方法降低40.79%(32M vs 54.1M),并引入后声码器对齐方法,使频谱失真(LSD)提升2.1倍,语音可懂度(STOI)提升4.7倍,信号失真比(SI-SDR)提升1.25倍。计算量减少13.33%的同时保持更高性能。据我们所知,首次展示如何通过细微的非线性混沌物理规律,结合Samba注意力机制,实现更小模型与更优性能的肌电转语音合成。

原文摘要 · Abstract (English)

We propose a chaos-inspired new architecture for EMG-to-Speech (ETS) synthesis called CS-ETS, which combines a Samba-based encoder with two novel chaos-inspired loss functions -- Lyapunov Exponent Regularization (LER) and Multi-Scale Detrended Fluctuation Analysis (MSDFA). LER is designed based on Lyapunov exponents to capture nonlinear fluctuations and sensitivity to initial conditions. MSDFA exploits detrended fluctuation analysis to quantify fractal-like, long-range temporal chaotic correlation. CS-ETS surpasses prior work with a 40.79\% lower parameter count (32M vs 54.1M) and introduces a new Post-Vocoder Alignment approach that improves LSD by 2.1x, STOI by 4.7x, and SI-SDR by 1.25x. CS-ETS reduces computation by 13.33\% while maintaining improved performance. To the best of our knowledge, for the first time, we show how ETS can be supervised by the subtle non-linear chaotic physics with Samba attention to achieve a significantly smaller model with superior performance.

肌电语音混沌建模小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。