arXiv:2607.21180cs.AI2026-07

为汽车语音助手设计安全防护机制,发现现有方案存在延迟和不可靠问题。

Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications

论文配图:Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications
图 1 · 摘自论文原文
  • 基于转录文本和工具调用两种安全防护方法
  • 检查每条回复延迟达0至1.4秒,影响实时交互
  • 适合关注车载语音系统安全性的工程师参考

近期进展带来了能生成自然对话(包括语调、情绪等非语言线索)的端到端语音到语音(S2S)对话助手。在汽车领域,这可实现直观且类人化的车内对话体验。然而,集成此类端到端助手会限制可编程的领域专用安全机制。本文探讨了两种S2S防护实现方式:基于转录文本与基于工具调用。通过实证评估,我们发现这两种策略在大多数工业部署场景下均不充分,主要因检查导致显著延迟(即使计算开销小,每条回复延迟仍达0至1.4秒),以及技术障碍(如工具调用行为可能非确定性)。最后,本文指出汽车场景中S2S防护面临的关键开放挑战。

原文摘要 · Abstract (English)

Recent advances have introduced speech-to-speech (S2S) conversational assistants capable of producing natural-sounding interactions, including non-verbal cues like tonality and mood. In the automotive domain, this enables intuitive and humanlike in-car dialogue experiences. However, integrating these end-to-end assistants limits architectural options for programmable domain-specific safeguards. This paper discusses two implementation approaches for S2S guardrails: transcript-based and tool-based. Through an empirical evaluation, we demonstrate that both strategies are insufficient for industrial deployment in most cases due to prohibitive latency (delaying each answer by 0 to 1.4 seconds even for computationally cheap checks) and technical impediments (like potentially non-deterministic tool call behavior). Finally, we outline open challenges for S2S guardrails in the automotive context.

语音助手安全防护汽车应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。