arXiv:2505.18516eess.AS2025-05

用可变长度的声学特征令牌替代固定帧,更精准捕捉抑郁检测中的时间动态。

Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection

  • 基于语言学理论动态分割语音,生成可变长度声学令牌
  • 在ADHD-10数据集上达到89.2%的抑郁检测准确率
  • 适合需要可解释性与时间敏感性的临床辅助诊断场景

大型语言模型(LLMs)在多个领域表现卓越,推动了将语音融入其框架的尝试,通常通过神经语音编解码器对连续音频进行分词,实现强大的语音语言模型。然而,主流分词策略依赖于固定时间间隔的均匀帧处理,这种固定速率方法虽适用于语言内容,却破坏了关键的时间动态信息——这些动态在临床应用中已被确立为抑郁症检测的核心生物标志物。为此,我们提出具有自适应特性的独特特征编解码器(Distinctive Feature Codec, DFC),旨在保留这一重要时间结构。借鉴语言学理论,DFC摒弃固定间隔处理,转而学习在感知显著的声学转换点处动态分割信号,生成能高效编码时间结构的可变长度令牌。作为主要贡献,本工作首次将传统独特特征整合进现代深度学习编解码器,用于时间敏感的抑郁症检测任务。同时引入分组标量量化(GSQ)方法,以稳定量化这些可变长度段落。基于独特特征的方法为传统帧基处理提供了有前景的替代方案,并推进了现代深度学习语音抑郁检测框架中的可解释表征学习。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable success across diverse fields, establishing a powerful paradigm for complex information processing. This has inspired the integration of speech into LLM frameworks, often by tokenizing continuous audio via neural speech codecs, enabling powerful speech language models. However, this dominant tokenization strategy relies on uniform frame-based processing at fixed time intervals. This fixed-rate approach, while effective for linguistic content, destroys the temporal dynamics. These dynamics are not noise but are established as primary biomarkers in clinical applications such as depression detection. To address this gap, we introduce the Distinctive Feature Codec (DFC), an adaptive framework engineered to preserve this vital timing information. Drawing from linguistic theory, DFC abandons fixed-interval processing and instead learns to dynamically segment the signal at perceptually significant acoustic transitions. This generates variable-length tokens that efficiently encode the temporal structure. As a key contribution, this work is the first to integrate traditional distinctive features into a modern deep learning codec for a temporally sensitive task such as depression detection. We also introduce the Group-wise Scalar Quantization (GSQ) approach to stably quantize these variable-length segments. Our distinctive feature-based approach offers a promising alternative to conventional frame-based processing and advances interpretable representation learning in the modern deep learning speech depression detection framework.

语音分析抑郁检测编解码器可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。