让语音理解模型彻底遗忘特定功能,防止被恶意诱导复现。
Selective Capability Unlearning in End-to-End Spoken Language Understanding

- 在表示层分离意图相关方向,抑制模型对指定意图的响应
- 在多个基准上使强制前缀恢复原意图的概率大幅下降
- 适合需要安全去功能化的语音系统部署场景
现代语音理解(SLU)系统在实际应用中常需因政策或安全原因移除特定功能,这些功能对应特定意图及其槽位生成行为。然而,在自回归模型中,即使抑制目标意图,其条件映射关系仍存在,当外部提供意图前缀时,模型可重建原始意图-槽位结构。我们将其称为‘能力持续性’问题。为此提出 extit{B}inding extit{S}ubspace (BSU) 框架,通过表示层隔离并衰减该映射所依赖的意图条件方向。在多个SLU基准上,BSU显著降低强制前缀下的意图恢复率,同时保持其余功能性能不变。
原文摘要 · Abstract (English)
Modern spoken language understanding (SLU) systems are increasingly deployed in real-world settings, where specific functionalities may need to be removed due to policy or safety constraints. In SLU, a functionality corresponds to an intent and its associated slot-generation behavior. However, in autoregressive models, suppressing a target intent does not eliminate the conditional mapping that generates slots conditioned on that intent. When the intent prefix is externally supplied, the model can reconstruct the original intent-slot structure. We identify this structural failure as \textbf{\emph{capability persistence}}. We propose \textit{\underline{B}inding \underline{S}ubspace (BSU)}, a representation-level framework that isolates and attenuates intent-conditioned directions underlying this mapping. Across SLU benchmarks, BSU substantially reduces forced-prefix recoverability while preserving retained performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。