TSUBASA通过动态记忆与自学习提升长时个性化,突破效率与质量瓶颈。
TSUBASA: Improving Long-Horizon Personalization via Evolving Memory and Self-Learning with Context Distillation
- 引入动态记忆演化与上下文蒸馏的自学习机制
- 在多尺寸Qwen模型上实现超越Mem0等系统的长时记忆表现
- 兼顾高效性与高保真度,适合长期个性化应用
个性化大语言模型(PLLM)虽能匹配用户需求,但在追踪长期对话或行为历史方面仍存短板。现有记忆机制难以捕捉动态变化,RAG面临质量与效率的权衡,参数化适配则受限于标注数据稀缺带来的训练-推理差距。为此,我们提出TSUBASA,采用双路径设计:通过动态记忆演化优化记忆写入,借助自学习与上下文蒸馏目标强化记忆读取,以内化用户经验。基于Qwen-3系列模型(4B至32B)在长时基准上的大量实验验证,TSUBASA显著优于依赖记忆写入的主流系统(如Mem0和Memory-R1)。分析表明,该方法打破质量-效率瓶颈,实现帕累托改进,在降低令牌开销的同时保持强鲁棒性与高保真个性化。
原文摘要 · Abstract (English)
Personalized large language models (PLLMs) have garnered significant attention for their ability to align outputs with individual's needs and preferences. However, they still struggle with long-horizon tasks, such as tracking a user's extensive history of conversations or activities. Existing memory mechanisms often fail to capture evolving behaviors, and RAG paradigms are trapped by a quality-efficiency tradeoff. Meanwhile, parametric adaptation is bottlenecked by train-inference gap due to the scarcity of labeled data. To enhance the long-horizon capabilities of PLLMs, we introduce TSUBASA, a two-pronged approach designed to improve memory writing via dynamic memory evolution, and memory reading via self-learning with a context distillation objective to internalize user experiences. Extensive evaluations on long-horizon benchmarks using the Qwen-3 model family (4B to 32B) validate the effectiveness of TSUBASA, surpassing competitive memory-augmented systems that rely primarily on memory writing, such as Mem0 and Memory-R1. Our analyses further confirms that TSUBASA breaks the quality-efficiency barrier to achieve Pareto improvements, delivering robust, high-fidelity personalization with a reduced token budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。