arXiv:2608.26005eess.AScs.AI2026-08

VoiceMem让对话系统有了实时、情感化、个性化的记忆能力。

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

论文配图:VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
图 1 · 摘自论文原文
  • 左右脑并行架构:左脑记事实,右脑管情绪与人格。
  • 检索仅需134毫秒,准确率比传统方法高30点。
  • 适合开发真实对话场景中的个性化智能助手。

对话系统如双工语音语言模型(SLMs)仍缺乏流式、精准且富有同理心的记忆系统。我们提出VoiceMem,一种包含并行信息左脑、情感右脑及流式内存输入输出机制的简单记忆架构。我们还构建了完整的训练管道,支持长时程评估和可替换后端的记忆解耦部署。实验与实际部署显示三大优势:(i) 准确性:在前5项召回下,左脑性能优于传统系统(如Mem0)在前200项召回下的表现,提升近30个百分点;(ii) 情感与个性化:右脑通过短/长时程情感归因与双节点人格建模,在三个角色基准上达到当前最优,综合得分提升4.29分;(iii) 实时且低成本:检索耗时仅134毫秒,低于标准语音活动检测延迟,不增加对话延迟,同时保持高准确率与低开销。结果表明,VoiceMem为实时、个性化、情感感知的语音交互提供了实用的记忆基础。

原文摘要 · Abstract (English)

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.

对话系统情感记忆实时交互双脑架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。