arXiv:2606.24063cs.CLcs.AI2026-06

让语音理解模型彻底遗忘特定功能,防止被恶意诱导复现。

Selective Capability Unlearning in End-to-End Spoken Language Understanding

论文配图:Selective Capability Unlearning in End-to-End Spoken Language Understanding
图 1 · 摘自论文原文
  • 在表示层分离意图相关方向,抑制模型对指定意图的响应
  • 在多个基准上使强制前缀恢复原意图的概率大幅下降
  • 适合需要安全去功能化的语音系统部署场景

现代语音理解(SLU)系统在实际应用中常需因政策或安全原因移除特定功能,这些功能对应特定意图及其槽位生成行为。然而,在自回归模型中,即使抑制目标意图,其条件映射关系仍存在,当外部提供意图前缀时,模型可重建原始意图-槽位结构。我们将其称为‘能力持续性’问题。为此提出 extit{​B}inding extit{​S}ubspace (BSU) 框架,通过表示层隔离并衰减该映射所依赖的意图条件方向。在多个SLU基准上,BSU显著降低强制前缀下的意图恢复率,同时保持其余功能性能不变。

原文摘要 · Abstract (English)

Modern spoken language understanding (SLU) systems are increasingly deployed in real-world settings, where specific functionalities may need to be removed due to policy or safety constraints. In SLU, a functionality corresponds to an intent and its associated slot-generation behavior. However, in autoregressive models, suppressing a target intent does not eliminate the conditional mapping that generates slots conditioned on that intent. When the intent prefix is externally supplied, the model can reconstruct the original intent-slot structure. We identify this structural failure as \textbf{\emph{capability persistence}}. We propose \textit{\underline{B}inding \underline{S}ubspace (BSU)}, a representation-level framework that isolates and attenuates intent-conditioned directions underlying this mapping. Across SLU benchmarks, BSU substantially reduces forced-prefix recoverability while preserving retained performance.

语音理解去功能化安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。