轻量级语音认证系统实现实时安全防护,无需额外硬件。
A Lightweight Dual-Factor Acoustic Authentication System via Cascaded GMM-DTW Architecture for Edge Computing

- 级联GMM-DTW架构分步验证说话人与口令,降低资源消耗。
- 防重放攻击误接受率仅6.67%,合法用户误拒绝率16.67%。
- 适合低功耗边缘设备部署,延迟稳定在9.82ms以内。
本文提出一种面向资源受限边缘环境的轻量级、级联式GMM-DTW双因子语音锁系统。通过共享的MFCC特征空间,框架采用顺序防御机制,结合GMM说话人筛选与DTW口令验证。为应对演示攻击且无需额外硬件,引入动态联合绝对-相对边界约束,使物理冒充和高保真重放攻击的误接受率(FAR)分别降至2.73%和6.67%,合法用户误拒绝率(FRR)为16.67%。借助Sakoe-Chiba窗口优化,单核CPU下端到端处理延迟在时序压力下严格控制在9.82ms内,包括1.51ms特征提取、0.54ms GMM打分和7.77ms最坏情况下的DTW匹配。实证基准表明,白盒声学级联方案可在低功耗边缘节点上实现安全、确定性的实时部署。
原文摘要 · Abstract (English)
This paper presents a lightweight, cascaded GMM-DTW dual-factor voice lock system for resource-constrained edge environments. By utilizing a shared MFCC feature space, the framework implements a sequential defense mechanism combining GMM speaker screening and DTW passphrase verification. To counter presentation threats without extra hardware, a dynamic joint absolute-relative margin constraint is integrated into the GMM classification space, limiting the physical imposter and high-fidelity replay attack False Acceptance Rates (FAR) to 2.73% and 6.67%, respectively, with a legitimate False Rejection Rate (FRR) of 16.67%. Due to Sakoe-Chiba window optimization, the global end-to-end processing latency under temporal stress is rigidly bounded at 9.82ms on a single-core CPU, comprising 1.51ms for feature extraction, 0.54ms for GMM scoring, and 7.77ms for worst-case DTW matching. These empirical benchmarks demonstrate the viability of white-box acoustic cascades for secure, deterministic real-time deployment on low-power edge nodes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。