arXiv:2607.01345cs.CLcs.AI2026-07被引 1

用概率模型自动评估对话中抢话、沉默等自然度问题

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

论文配图:TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
图 1 · 摘自论文原文
  • 基于因果预测模型计算说话人未来语音活动的似然值
  • 通过负对数似然衡量对话时序异常程度,识别不自然交互
  • 适用于评估多类型对话失误,适合对话系统研发者使用

对话中的自然轮流机制是全双工语音对话系统的核心,但其自动化评估仍受限。现有方法多依赖人工评分或特定时间指标,难以在统一框架下比较不同类型的时序异常。本文提出TurnNat,一种基于似然的双人对话轮流自然度自动评估框架。该框架训练因果轮流预测模型以估计未来双说话人的语音活动状态,利用观测到的未来活动的负对数似然(NLL)衡量时序异常程度。TurnNat将帧级NLL在从话语起止点提取的轮流边界单元(TBUs)上聚合,并将均值与尾部TBU得分合并为对话级自然度分数。我们构建了一个包含成对自然与扰动对话片段的受控扰动基准,经人工自然度判断验证。实验表明,TurnNat能有效识别多种异构时序异常下的不自然轮流行为。

原文摘要 · Abstract (English)

Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluations often rely on human judgments or behavior-specific timing metrics, making it difficult to compare heterogeneous timing failures within a unified framework. We propose TurnNat, a likelihood-based framework for automatic turn-taking naturalness evaluation in two-channel spoken dialogue. A causal turn-taking prediction model trained on natural conversations estimates future two-speaker voice-activity states, and the negative log-likelihood (NLL) of the observed future activity measures timing atypicality. TurnNat pools frame-level NLLs over turn-taking boundary units (TBUs) extracted from utterance onsets and offsets, and aggregates mean and tail TBU scores into a dialogue-level naturalness score. We further construct a controlled perturbation benchmark of paired natural and perturbed dialogue clips, validated by human naturalness judgments. Experiments on this benchmark show that TurnNat successfully identifies unnatural turn-taking perturbations across heterogeneous timing failures.

对话系统自然度评估语音活动检测轮流机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。