arXiv:2509.13989eess.AS2025-09中稿 · ICASSP 2026被引 1

评测语音指令系统中用户指令与听觉感知的差距,发现多数模型对情感和年龄控制不准。

Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems

  • 通过真人评分分析指令语音在强度和情绪上的表现力控制效果。
  • gpt-4o-mini-tts 最接近指令意图,但多数模型误判年龄属性。
  • 细粒度控制仍不足,适合语音合成与人机交互研究者参考。

指令引导式文本转语音(ITTS)让用户通过自然语言控制语音生成,提供更直观的交互方式。然而,用户风格指令与听众感知之间的对齐程度尚未深入探究。本文首次对两种表达维度(程度副词和情绪强度分级)进行感知分析,并收集了关于说话人年龄与词级强调属性的人类评分。为全面揭示指令-感知差距,构建了大规模人类评估数据集,命名为表达性语音控制(E-VOC)语料库。结果表明:(1)gpt-4o-mini-tts 是最可靠的 ITTS 模型,在声学维度上指令与生成语音的对齐度最佳;(2)所分析的5个 ITTS 系统普遍生成成人声音,即使指令要求儿童或老年声音;(3)细粒度控制仍是主要挑战,说明多数 ITTS 系统在理解细微属性指令方面仍有巨大改进空间。

原文摘要 · Abstract (English)

Instruction-guided text-to-speech (ITTS) enables users to control speech generation through natural language prompts, offering a more intuitive interface than traditional TTS. However, the alignment between user style instructions and listener perception remains largely unexplored. This work first presents a perceptual analysis of ITTS controllability across two expressive dimensions (adverbs of degree and graded emotion intensity) and collects human ratings on speaker age and word-level emphasis attributes. To comprehensively reveal the instruction-perception gap, we provide a data collection with large-scale human evaluations, named Expressive VOice Control (E-VOC) corpus. Furthermore, we reveal that (1) gpt-4o-mini-tts is the most reliable ITTS model with great alignment between instruction and generated utterances across acoustic dimensions. (2) The 5 analyzed ITTS systems tend to generate Adult voices even when the instructions ask to use child or Elderly voices. (3) Fine-grained control remains a major challenge, indicating that most ITTS systems have substantial room for improvement in interpreting slightly different attribute instructions.

语音合成指令控制感知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。