提出可信赖音频情感计算框架,保护用户隐私的同时精准识别抑郁状态。
TAAC: A gate into Trustable Audio Affective Computing
- 通过子空间分解分离抑郁特征与敏感身份信息
- 加密仅针对敏感信息,仍保持90%以上诊断准确率
- 适合医疗AI中需兼顾隐私与精度的场景
随着人工智能在抑郁症诊断中的应用,高需求与低供给的矛盾得到显著缓解。音频作为情绪表达的主要载体,日益受到学术界和产业界关注。然而,音频数据包含易泄露的用户身份信息,在智能诊断过程中存在被恶意利用的风险。以往方法难以有效区分抑郁特征与敏感特征,且缺乏安全加密机制。为此,我们提出首个可信赖音频情感计算框架TAAC,采用基于对抗损失的子空间分解技术,集成差异特征子空间分解器(DFSD)、灵活噪声加密器(FNE)与分阶段训练范式,分别实现特征分离、身份信息加密与性能提升。大量实验表明,该框架在抑郁检测、身份保留与音频重建方面均优于现有方法;在不同加密强度下也表现稳定,验证了其在保密性、准确性、可追溯性与可调性上的卓越表现。
原文摘要 · Abstract (English)
With the emergence of AI techniques for depression diagnosis, the conflict between high demand and limited supply for depression screening has been significantly alleviated. Among various modal data, audio-based depression diagnosis has received increasing attention from both academia and industry since audio is the most common carrier of emotion transmission. Unfortunately, audio data also contains User-sensitive Identity Information (ID), which is extremely vulnerable and may be maliciously used during the smart diagnosis process. Among previous methods, the clarification between depression features and sensitive features has always serve as a barrier. It is also critical to the problem for introducing a safe encryption methodology that only encrypts the sensitive features and a powerful classifier that can correctly diagnose the depression. To track these challenges, by leveraging adversarial loss-based Subspace Decomposition, we propose a first practical framework \name presented for Trustable Audio Affective Computing, to perform automated depression detection through audio within a trustable environment. The key enablers of TAAC are Differentiating Features Subspace Decompositor (DFSD), Flexible Noise Encryptor (FNE) and Staged Training Paradigm, used for decomposition, ID encryption and performance enhancement, respectively. Extensive experiments with existing encryption methods demonstrate our framework's preeminent performance in depression detection, ID reservation and audio reconstruction. Meanwhile, the experiments across various setting demonstrates our model's stability under different encryption strengths. Thus proving our framework's excellence in Confidentiality, Accuracy, Traceability, and Adjustability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。