arXiv:2607.05901cs.AI2026-07

通过加权排序优化抑郁检测的隐空间分布,提升判别能力。

Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking

论文配图:Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking
图 1 · 摘自论文原文
  • 设计双机制损失函数,动态加权难样本对并压缩类内差异。
  • 在D-vlog和LMVD数据集上达到当前最佳性能。
  • 适合关注多模态情感识别与隐空间建模的研究者。

基于音视频数据的自动抑郁检测面临特征分布重叠和决策边界不稳健等挑战。为此,我们提出一种细粒度多模态框架,包含时序编码器与互注意力变压器,实现深层跨模态融合。核心贡献是二分类优势加权排序损失(Binary Advantage-weighting Ranking Loss),通过两个互补机制优化隐空间分布:优势加权分离机制通过计算成对预测差矩阵,动态加权困难样本对;优势加权紧凑性机制则最小化类内方差,促使特征聚类于各自类别中心。在D-vlog和LMVD数据集上的大量实验表明,该模型通过优先处理困难样本重建了隐空间的序数结构,从而实现领先性能。

原文摘要 · Abstract (English)

Automatic depression detection using audio-visual data faces significant challenges, particularly in disentangling overlapping feature distributions and establishing robust decision boundaries. To address this, we propose a fine-grained multimodal framework featuring a temporal encoder and a mutual transformer to facilitate deep cross-modal fusion. Our core contribution is the Binary Advantage-weighting Ranking Loss, which optimizes the latent space distribution through two complementary mechanisms: Advantage-weighted Separation, which mines hard pairs by computing a pairwise prediction difference matrix and dynamically weighting them based on their difficulty; and Advantage-weighted Compactness, which minimizes intra-class variance to force features to cluster around their respective class centers. Extensive experiments on D-vlog and LMVD demonstrate that our model reconstructs the latent ordinal structure by prioritizing hard pairs, thereby achieving state-of-the-art performance.

抑郁检测多模态隐空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。