arXiv:2509.04821cs.CL2025-09中稿 · IEEE ICASSP 2026被引 1

用自适应特征蒸馏提升语音理解模型性能,解决数据少与部署难问题

AFD-SLU: Adaptive Feature Distillation for Spoken Language Understanding

  • 通过动态适配器对齐教师与学生模型的异构特征空间
  • 在中文口语理解基准上达到95.67%意图准确率和92.02%槽位F1
  • 适合资源受限场景下高效部署高精度语音理解系统

语音语言理解(SLU)是对话系统的核心,使机器能解析用户语句。然而,由于标注数据稀缺以及大型语言模型(LLMs)在实际应用中计算开销大,构建高效SLU系统仍具挑战。为此,我们提出自适应特征蒸馏框架(AFD-SLU),将基于通用文本嵌入(GTE)的教师模型中的丰富语义表示迁移到轻量级学生模型。方法引入配备残差投影神经网络(RPNN)的动态适配器以对齐异构特征空间,并设计动态蒸馏系数(DDC),根据意图和槽位预测性能实时调节蒸馏强度。在基于中文的ProSLU基准测试中,AFD-SLU取得当前最优效果:意图准确率达95.67%,槽位F1为92.02%,整体准确率为85.50%。

原文摘要 · Abstract (English)

Spoken Language Understanding (SLU) is a core component of conversational systems, enabling machines to interpret user utterances. Despite its importance, developing effective SLU systems remains challenging due to the scarcity of labeled training data and the computational burden of deploying Large Language Models (LLMs) in real-world applications. To further alleviate these issues, we propose an Adaptive Feature Distillation framework that transfers rich semantic representations from a General Text Embeddings (GTE)-based teacher model to a lightweight student model. Our method introduces a dynamic adapter equipped with a Residual Projection Neural Network (RPNN) to align heterogeneous feature spaces, and a Dynamic Distillation Coefficient (DDC) that adaptively modulates the distillation strength based on real-time feedback from intent and slot prediction performance. Experiments on the Chinese profile-based ProSLU benchmark demonstrate that AFD-SLU achieves state-of-the-art results, with 95.67% intent accuracy, 92.02% slot F1 score, and 85.50% overall accuracy.

语音理解特征蒸馏轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。