arXiv:2608.04351cs.SDcs.AI2026-08

用双曲几何提升语音情感微调效率,更适配情绪多层次特征。

HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models

论文配图:HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models
图 1 · 摘自论文原文
  • 在双曲空间中建模情绪粒度,用半径显式表示从语调到语义的多层级特征。
  • 在MELD上全面超越欧氏基线,在IEMOCAP上显著提升少数类识别准确率。
  • 适合资源受限下需要高精度情绪识别的语音大模型微调场景。

大型音频-语言模型(LALMs)在通用语音理解上表现优异,但将其适配至细粒度任务如语音情感识别(SER)仍存在显著瓶颈。现有参数高效微调(PEFT)方法通常在平坦的欧几里得空间中操作,该几何结构难以捕捉情绪线索的多粒度特性——从低层次韵律到高层次语义。为此,我们提出HyPASE,一种基于双曲几何的LALM-SER参数高效微调框架。HyPASE采用庞加莱球模型,以双曲半径作为表征粒度的显式代理。框架包含两个核心组件:用于层自适应权重调制的双曲几何适配器(HGA),以及将多尺度特征压缩为紧凑音频前缀的情感感知多容量跨模态聚合器(EMCA)。在标准基准上的实证结果表明,HyPASE在所有指标上均优于欧几里得基线,于MELD上表现全面领先;在IEMOCAP上实现显著的未加权准确率提升,尤其在类别不平衡的情感识别中表现突出,伴随轻微加权准确率下降,反映出双曲空间对少数类表征的几何优先性;此外,HyPASE在有限参数预算下实现了稳健的零样本跨数据集泛化能力。通过将适配过程建立在双曲几何基础上,HyPASE为LALM微调提供了一条高效路径。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion Recognition (SER) remains a significant bottleneck. Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate in flat Euclidean space, and this geometry fails to capture the multi-granularity nature of emotion cues, which range from low-level prosody to high-level semantics. To address this, we propose HyPASE, a hyperbolic PEFT framework for LALM-based SER. HyPASE leverages the Poincare ball model, using the hyperbolic radius as an explicit proxy for representational granularity. The framework consists of two core components: a Hyperbolic Geometric Adapter (HGA) for layer-adaptive weight modulation, and an Emotion-aware Multi-capacity Cross-modal Aggregator (EMCA) that compresses multi-scale features into compact audio prefixes. Empirical results on standard benchmarks show that HyPASE outperforms Euclidean PEFT baselines across all metrics on MELD and achieves a notable Unweighted Accuracy gain on IEMOCAP, particularly in class-imbalanced emotion recognition, with the accompanying slight Weighted Accuracy trade-off reflecting hyperbolic space's geometric prioritization of minority-class representations; furthermore, HyPASE achieves robust zero-shot cross-dataset generalization within a constrained parameter budget. By grounding the adaptation process in hyperbolic geometry, HyPASE offers a highly efficient path for LALM fine-tuning.

语音情感双曲几何参数高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。