arXiv:2607.23204cs.ROcs.CL2026-07中稿 · ICMI LBR 2026

用智能预读减少对话延迟,让机器人更自然地接话

Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot

论文配图:Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot
图 1 · 摘自论文原文
  • 先预测用户意图,提前生成简短回应
  • 实测显示预读能缩短主回复等待时间
  • 适合需要实时交互的智能客服场景

基于大语言模型的对话系统因需等语音识别完成才开始生成,存在响应延迟。现有固定填充词虽可缓解,但使用久了显得不自然。本文提出一种两阶段增量框架,将预回复准备与语音启动解耦:当用户意图可预测时,意图就绪检测器触发大模型生成简短预回复;同时,语音活动预测(VAP)模型决定何时输出。在商场路线引导机器人的真实场景实验中,比较了无填充、固定填充和上下文预读三种条件。结果表明,固定填充和上下文预读均显著降低初始响应延迟,相比固定填充,上下文预读初始延迟略高,但初始到主回复的间隔显著缩短。探索性评分未见显著差异,说明存在时间权衡。

原文摘要 · Abstract (English)

Large language model (LLM)-based dialogue systems suffer response delays because generation begins only after final speech recognition. While fixed fillers are a workaround, they become unnatural over time. We propose a two-stage incremental framework that decouples prefatory-response preparation from speech onset. Once user intent becomes predictable, an intent readiness detector triggers LLM-based generation of a short prefatory response. Concurrently, a voice activity projection (VAP) model determines when to deliver it. Through a field experiment with a route-guidance robot in a shopping mall, we evaluated three conditions: no-filler, fixed-filler, and contextual-preface. Both fixed-filler and contextual-preface significantly reduced initial response latency relative to no-filler. Relative to fixed-filler, contextual-preface had significantly longer initial response latency but a significantly shorter initial-to-main gap. Exploratory ratings showed no significant differences. These results indicate a timing trade-off.

对话系统低延迟预读生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。