arXiv:2501.09645cs.AIcs.CL2025-01中稿 · presentation at th…被引 10

用分类框架提升语音助手长期记忆,兼顾个性化与隐私安全。

CarMem: Enhancing Long-Term Memory in LLM Voice Assistants through Category-Bounding

  • 以预定义类别构建记忆体系,利用大模型高效提取存储偏好。
  • 在多轮多会话数据集上,偏好提取F1达0.78至0.95,检索准确率达0.87。
  • 减少95%冗余、92%矛盾偏好,适合注重隐私的工业级语音系统。

当前语音助手虽能提升交互体验与用户粘性,但普遍难以持久保留用户偏好,导致重复请求与用户流失。此外,行业应用中缺乏监管的偏好提取方式引发隐私与信任危机,尤其在欧洲等严苛监管地区。为此,我们提出一种基于预设类别的长期记忆系统,利用大语言模型实现偏好在分类框架内的高效提取、存储与检索,兼顾个性化与透明性。我们构建了一个基于真实工业数据的合成多轮多会话对话数据集CarMem,专用于车载语音助手场景。在该数据集上,系统在不同类别粒度下偏好提取的F1得分介于0.78至0.95之间;维护策略使冗余偏好减少95%,矛盾偏好减少92%;最优检索准确率为0.87。结果表明该系统具备工业应用潜力。

原文摘要 · Abstract (English)

In today's assistant landscape, personalisation enhances interactions, fosters long-term relationships, and deepens engagement. However, many systems struggle with retaining user preferences, leading to repetitive user requests and disengagement. Furthermore, the unregulated and opaque extraction of user preferences in industry applications raises significant concerns about privacy and trust, especially in regions with stringent regulations like Europe. In response to these challenges, we propose a long-term memory system for voice assistants, structured around predefined categories. This approach leverages Large Language Models to efficiently extract, store, and retrieve preferences within these categories, ensuring both personalisation and transparency. We also introduce a synthetic multi-turn, multi-session conversation dataset (CarMem), grounded in real industry data, tailored to an in-car voice assistant setting. Benchmarked on the dataset, our system achieves an F1-score of .78 to .95 in preference extraction, depending on category granularity. Our maintenance strategy reduces redundant preferences by 95% and contradictory ones by 92%, while the accuracy of optimal retrieval is at .87. Collectively, the results demonstrate the system's suitability for industrial applications.

语音助手长期记忆隐私保护大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。