评测开源语音模型记忆与复用对话上下文的能力,发现其表现明显弱于闭源模型。
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models

- 构建新基准ContextDialog,系统评估开源模型对历史对话的利用能力。
- 语音模型在回忆语音内容时表现更差,即使使用检索增强生成仍难应对相关问题。
- 适合关注语音助手记忆能力的研究者与开发者参考。
近期多轮语音交互模型提升了用户与模型的沟通效率。然而,尽管闭源模型能有效保留并召回历史语句,开源模型是否具备相同能力尚不明确。为填补这一空白,我们提出ContextDialog基准,系统评估开源交互模型对过去语句的利用程度。结果表明,基于语音的模型在回忆语音信息方面比文本模型更困难,即使采用检索增强生成,模型在回答涉及历史对话的问题时仍表现不佳。这些发现揭示了开源模型在记忆保持与检索鲁棒性方面的关键局限,并为改进提供方向。
原文摘要 · Abstract (English)
Recent advancements in multi-turn voice interaction models have improved user-model communication. However, while closed-source models effectively retain and recall past utterances, whether open-source models share this ability remains unexplored. To fill this gap, we systematically evaluate how well open-source interaction models utilize past utterances using ContextDialog, a benchmark we proposed for this purpose. Our findings show that speech-based models have more difficulty than text-based ones, especially when recalling information conveyed in speech, and even with retrieval-augmented generation, models still struggle with questions about past utterances. These insights highlight key limitations in open-source models and suggest ways to improve memory retention and retrieval robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。