arXiv:2502.19759cs.SDeess.AS2025-02ACL被引 7

评测开源语音模型记忆与复用对话上下文的能力,发现其表现明显弱于闭源模型。

Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models

论文配图:Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
图 1 · 摘自论文原文
  • 构建新基准ContextDialog,系统评估开源模型对历史对话的利用能力。
  • 语音模型在回忆语音内容时表现更差,即使使用检索增强生成仍难应对相关问题。
  • 适合关注语音助手记忆能力的研究者与开发者参考。

近期多轮语音交互模型提升了用户与模型的沟通效率。然而,尽管闭源模型能有效保留并召回历史语句,开源模型是否具备相同能力尚不明确。为填补这一空白,我们提出ContextDialog基准,系统评估开源交互模型对过去语句的利用程度。结果表明,基于语音的模型在回忆语音信息方面比文本模型更困难,即使采用检索增强生成,模型在回答涉及历史对话的问题时仍表现不佳。这些发现揭示了开源模型在记忆保持与检索鲁棒性方面的关键局限,并为改进提供方向。

原文摘要 · Abstract (English)

Recent advancements in multi-turn voice interaction models have improved user-model communication. However, while closed-source models effectively retain and recall past utterances, whether open-source models share this ability remains unexplored. To fill this gap, we systematically evaluate how well open-source interaction models utilize past utterances using ContextDialog, a benchmark we proposed for this purpose. Our findings show that speech-based models have more difficulty than text-based ones, especially when recalling information conveyed in speech, and even with retrieval-augmented generation, models still struggle with questions about past utterances. These insights highlight key limitations in open-source models and suggest ways to improve memory retention and retrieval robustness.

语音交互对话记忆开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。