arXiv:2604.10065cs.CLcs.AI2026-04被引 3

分离说话时机与内容生成,提升语音对话模型的自然互动性。

ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models

论文配图:ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
图 1 · 摘自论文原文
  • 将说话时机与文本内容解耦,用二值状态控制发言与否。
  • 相比标准GRPO,重复n-gram比例降低超50%,有效抑制生成崩溃。
  • 适合需要精准轮次控制的实时语音交互系统研发者。

端到端全双工语音语言模型(SLMs)需要精确的轮次切换以实现自然交互。然而,通过标准原始词元强化学习(RL)优化时间动态会损害语义质量,导致严重生成崩溃和重复。我们提出ASPIRin,一种面向交互优化的强化学习框架,明确解耦何时说话与说什么。通过动作空间投影,将文本词汇映射为粗粒度二值状态(活跃发言或静默)。结合基于规则的奖励机制,使用组相对策略优化(GRPO),平衡用户打断与响应延迟。实证评估表明,ASPIRin在轮次切换、回应性短语及停顿处理方面均优化了交互性。关键在于,将时间与词元选择隔离,保持了语义连贯性,相较于标准GRPO,重复n-gram占比降低超过50%,有效消除退化性重复。

原文摘要 · Abstract (English)

End-to-end full-duplex Speech Language Models (SLMs) require precise turn-taking for natural interaction. However, optimizing temporal dynamics via standard raw-token reinforcement learning (RL) degrades semantic quality, causing severe generative collapse and repetition. We propose ASPIRin, an interactivity-optimized RL framework that explicitly decouples when to speak from what to say. Using Action Space Projection, ASPIRin maps the text vocabulary into a coarse-grained binary state (active speech vs. inactive silence). By applying Group Relative Policy Optimization (GRPO) with rule-based rewards, it balances user interruption and response latency. Empirical evaluations show ASPIRin optimizes interactivity across turn-taking, backchanneling, and pause handling. Crucially, isolating timing from token selection preserves semantic coherence and reduces the portion of duplicate n-grams by over 50% compared to standard GRPO, effectively eliminating degenerative repetition.

语音交互强化学习生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。