arXiv:2605.03039cs.LGcs.AI2026-05

用精度差异分离语音中的说话人特征与情绪状态,适合边缘设备实时检测躁狂

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection

论文配图:Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection
图 1 · 摘自论文原文
  • 用不同精度量化编码特征与状态,形成信息瓶颈实现无对抗分离
  • 在833人数据上达0.117相关性,优于现有方法2.8至15.9点
  • 仅617KB内存占用,23.4毫秒延迟,可部署于低价设备

通过语音生物标志物连续监测双相障碍躁狂状态,需在资源受限的边缘设备上分离稳定的说话人特征与易变的情绪状态。本文提出MP-IB框架,首次将混合精度量化作为临床特征-状态分离的信息瓶颈。核心思想是数值精度本身控制容量:FP16特征头(1,024位)编码说话人身份,INT4状态头(128位)捕捉躁狂程度,实现8倍信息不对称,无需对抗训练。结合动态精度调度与多尺度时序融合。在Bridge2AI-Voice数据集(N=833,每人4次会话,严格说话人独立交叉验证)上,MP-IB达成rho=0.117(95% CI: [0.089, 0.145],p=0.003 vs. 随机),优于9400万参数的WavLM-Adapter(rho=-0.042)、beta VAE(rho=0.089)和人工设计韵律特征(rho=0.031),绝对提升2.8–15.9点。零样本迁移至CREMA-D数据集,AUC达0.817。身份泄露降至近随机水平(EER=0.42,MIA-AUC=0.52)。端到端延迟仅23.4毫秒,模型大小617 KB,可在20美元以下设备实现实时监测。

原文摘要 · Abstract (English)

Continuous monitoring of bipolar disorder agitation via voice biomarkers requires disentangling stable speaker traits from volatile affective states on resource-constrained edge devices. We introduce MP-IB, the first framework to treat mixed-precision quantization as an information bottleneck for clinical trait-state separation. The core insight is that numerical precision itself controls capacity: an FP16 trait head (1,024 bits) encodes speaker identity, while an INT4 state head (128 bits) captures agitation, yielding 8x information asymmetry without adversarial training. We augment this with Dynamic Precision Scheduling and Multi-Scale Temporal Fusion. On Bridge2AI-Voice (N=833, 4 sessions/participant, strict speaker-independent CV), MP-IB achieves rho = 0.117 (95\% CI: [0.089, 0.145], p=0.003 vs. chance), outperforming 94M-parameter WavLM-Adapter with in-domain SSL continuation (rho = -0.042), beta VAE disentanglement (rho = 0.089), and hand-crafted prosody (rho = 0.031) by 2.8--15.9 points absolute. Zero-shot transfer to CREMA-D achieves AUC=0.817. Identity leakage is suppressed to near-random (EER=0.42, MIA-AUC=0.52). End-to-end latency is 23.4 ms with a 617 KB footprint, enabling real-time monitoring on sub 20 dollar devices.

边缘计算语音分析情绪识别模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。