arXiv:2604.25719eess.AS2026-04被引 6

用人类反馈强化学习提升语音模型对话自然度,打破验证奖励陷阱。

Step-Audio-R1.5 Technical Report

论文配图:Step-Audio-R1.5 Technical Report
图 1 · 摘自论文原文
  • 改用人类反馈强化学习(RLHF),避免过度追求文本验证正确性
  • 在保持分析推理能力的同时显著提升语音对话的自然流畅度
  • 特别适合需要沉浸式长对话体验的语音交互场景

近期大型音频语言模型将思维链(CoT)推理拓展至听觉领域,推动模型处理更复杂的声学与口语任务。当前主流方法依赖基于验证奖励的强化学习(RLVR),但该方法迫使模型将连续的音频语境压缩为孤立的可验证文本标签,导致一个根本问题:我们是在培养真正的音频智能,还是只是将连续感官信息简化为离散谜题?我们称之为“可验证奖励陷阱”。尽管RLVR在标准化客观基准上表现优异,却系统性损害了音频模型的真实对话感受。因过度强调孤立正确性,忽视声调自然度、情感连贯性与用户沉浸感,尤其在长轮对话中表现退化。为此,我们提出Step-Audio-R1.5,标志着向音频推理中人类反馈强化学习(RLHF)的根本转变。综合评估表明,Step-Audio-R1.5不仅维持了强大的分析推理能力,更显著重塑交互体验,重新定义深度沉浸式长轮口语对话的边界。

原文摘要 · Abstract (English)

Recent advancements in large audio language models have extended Chain-of-Thought (CoT) reasoning into the auditory domain, enabling models to tackle increasingly complex acoustic and spoken tasks. To elicit and sustain these extended reasoning chains, the prevailing paradigm -- driven by the success of text-based reasoning models -- overwhelmingly relies on Reinforcement Learning with Verified Rewards (RLVR). However, as models are strictly optimized to distill rich, continuous auditory contexts into isolated, verifiable text labels, a fundamental question arises: are we fostering true audio intelligence, or merely reducing a continuous sensory medium into a discrete puzzle? We identify this as the "verifiable reward trap." While RLVR yields remarkable scores on standardized objective benchmarks, it systematically degrades the real-world conversational feel of audio models. By prioritizing isolated correctness over acoustic nuance, RLVR reduces dynamic interactions to mechanical "answering machines," severely compromising prosodic naturalness, emotional continuity, and user immersion, particularly in long-turn dialogues. To bridge the gap between mechanical objective verification and genuine sensory empathy, we introduce Step-Audio-R1.5, marking a paradigm shift toward Reinforcement Learning from Human Feedback (RLHF) in audio reasoning. Comprehensive evaluations demonstrate that Step-Audio-R1.5 not only maintains robust analytical reasoning but profoundly transforms the interactive experience, redefining the boundaries of deeply immersive long-turn spoken dialogue.

语音生成强化学习对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。