arXiv:2609.03203cs.SDcs.CL2026-09

提出无需听者的语音规划评估方法,检验语音表达是否真实源自原文。

VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis

  • 通过证据引用与规则验证,检测语音规划是否基于原文内容
  • 在1440个案例中,未引用原文时规划准确率下降至0.684,证明来源依赖的重要性
  • 适合语音合成、对话系统等需严格语义对齐的场景

富有表现力的语音系统在生成波形前就已决定语调、情感、节奏等表达方式。在对话代理、叙述和角色驱动文本转语音中,这些隐藏的规划步骤会影响语调、音高、能量、语速、停顿、强调和立场,但下游音频评分往往无法揭示这些决策是否真正源自原始文本,这种源文本脱节的问题发生在任何波形生成之前。VoxReason将这一预合成阶段的决策可度量化,实现无需听者的源文本对齐语音规划评估。系统输出带证据引用的语音规划方案,由确定性验证器检查引用合法性、槽位一致性、无依据状态、模式有效性及单线索反事实局部性。在1440个经验证的源-标签案例中,快捷控制表明仅关注槽位准确率不安全:一个关键查找最优解达到1.000的槽位准确率,而情绪先验在无源键关联情况下仍保持0.958的准确率,且未引用强度或身份。在另100个学习型源键不匹配案例中,7B规模的局部性微调+反事实修复使槽位准确率/局部性从0.684/0.141提升至0.919/1.000,移除源文本后引用必要性得分下降0.488。生成波形质量不在当前评估范围内。

原文摘要 · Abstract (English)

Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In dialogue agents, narration, and role-conditioned TTS, that hidden planning step sets affect, pitch, energy, rate, pause, emphasis, and stance, yet downstream audio scores rarely reveal whether those choices were licensed by the source record, a source-use failure that occurs before any waveform exists. VoxReason makes that pre-synthesis decision measurable as a listener-free task for source-grounded speech planning. Before synthesis, VoxReason measures whether delivery choices are grounded in cited source records. Systems output a source-cited speaking-plan with evidence citations, and a deterministic verifier checks citation legality, slot agreement, unsupported state, schema validity, and one-cue counterfactual locality. On 1,440 checked source-label cases, shortcut controls show why slot accuracy alone is unsafe: a key-lookup oracle reaches 1.000 plan-slot accuracy on seen keys, while an emotion prior still reaches 0.958 slot accuracy on source-key-disjoint cases without citing intensity or identity. In a separate 100-case learned source-key-disjoint comparison, a 7B locality SFT+CF repair improves plan-slot accuracy/locality from 0.684/0.141 to 0.919/1.000, and removing source records lowers citation-required grounded score by 0.488. Rendered waveform quality remains outside the present evaluation.

语音合成源对齐评估方法自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。