arXiv:2512.11724cs.HCcs.AI2025-12

研究语音对话系统中模块化设计导致的对话断裂问题

From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines

  • 分析模块化语音转语音系统中的交互摩擦机制
  • 发现三类对话断裂模式:时间错位、表达扁平、修复僵化
  • 建议从组件优化转向接口协同设计,提升对话自然度

尽管基于语音的AI系统在生成能力上取得显著进展,其交互体验却常显不自然。本文研究模块化语音转语音检索增强生成(S2S-RAG)系统中的交互摩擦。通过对典型生产系统分析,超越简单的延迟指标,识别出三类反复出现的对话断裂现象:(1) 时间错位,系统延迟违背用户对对话节奏的预期;(2) 表达扁平化,副语言线索缺失导致回复过于字面化且不恰当;(3) 修复僵化,架构限制使用户无法实时纠正错误。系统级分析表明,这些摩擦点并非缺陷,而是模块化设计优先控制而非流畅性的结构性后果。结论指出,构建自然语音AI是基础设施设计挑战,需从优化单个组件转向精心协调各模块间的衔接。

原文摘要 · Abstract (English)

While voice-based AI systems have achieved remarkable generative capabilities, their interactions often feel conversationally broken. This paper examines the interactional friction that emerges in modular Speech-to-Speech Retrieval-Augmented Generation (S2S-RAG) pipelines. By analyzing a representative production system, we move beyond simple latency metrics to identify three recurring patterns of conversational breakdown: (1) Temporal Misalignment, where system delays violate user expectations of conversational rhythm; (2) Expressive Flattening, where the loss of paralinguistic cues leads to literal, inappropriate responses; and (3) Repair Rigidity, where architectural gating prevents users from correcting errors in real-time. Through system-level analysis, we demonstrate that these friction points should not be understood as defects or failures, but as structural consequences of a modular design that prioritizes control over fluidity. We conclude that building natural spoken AI is an infrastructure design challenge, requiring a shift from optimizing isolated components to carefully choreographing the seams between them.

语音交互模块化系统对话流畅性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。