arXiv:2608.19515cs.CL2026-08

测试语音语调如何影响助手决策,发现语音信息在文字不足时至关重要。

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

论文配图:Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does
图 1 · 摘自论文原文
  • 设计统一评测框架,对比文字与语音输入下的助手决策差异。
  • 语音隐含信息下模型决策准确率从14.6%提升至39.6%,文字明确表达时提升有限。
  • 语音能力模型需显式提取语音信息才能有效指导行动,否则难以转化。

语调线索可传递任务相关的信息,改变对话任务的走向与结果,即使词汇内容不变。现有基准通常孤立评估语调感知、回复恰当性与任务导向对话,难以检验语调证据是否影响下游决策。我们提出Hear2Act,一个统一的文本与语音助手评测协议,包含480个基于角色的情景,隐藏用户关切,且结果可客观验证。每个情景保持任务与用户需求不变,仅变化同一关切是通过文字明确表达还是主要通过语调传递,并在仅文本、音频及语境状态三种条件下评估决策。使用Hear2Act评估两个具备语音能力的大语言模型。在语调驱动反馈下,仅添加音频使最优解率从14.6%提升至15.3%;而当模型从音频中推断关切状态并以文本形式表示,再用于下一步动作选择时,准确率升至39.6%,接近40.7%的真值状态水平。但在文字明确表达关切时,该差异几乎消失。结果表明,当词汇证据不足时,语调具有关键作用,且语音大模型虽能从语音中恢复信息,但若无显式的中间表示,无法可靠地转化为行动。

原文摘要 · Abstract (English)

Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions. We introduce Hear2Act, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes. For each scenario, we keep the task and user needs fixed while varying whether the same concern is conveyed explicitly in words or primarily through prosody, and evaluate decisions under transcript, audio, and concern-state access. Using Hear2Act, we evaluate two audio-capable LLMs. Under Prosody-mediated feedback, adding audio to the transcript changes the average optimal-solution rate only from 14.6% to 15.3%. In contrast, when models infer the concern status from audio, represent it in text, and use it for next-action selection, the rate rises to 39.6%, close to 40.7% with the ground-truth state. This contrast, however, largely disappears under Explicit lexical feedback, where the concern is verbally mentioned in the utterance. Together, these results show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do not reliably carry it into action without an explicit intermediate representation.

语音理解多模态对话系统评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。