只在私人记忆能提升回答质量时才使用,避免无效信息干扰。
TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Personalized Generation

- 分两阶段:先查缺补漏,再按实用价值选择性引入记忆
- 在4500个任务中优于随机、词法和语义检索方法
- 适合需要精准个性化但怕冗余的生成场景
个性化生成系统通过召回用户历史来增强回复,但相关记忆可能偏离偏好、重复公共信息或支持不足。本文提出TRACE-Memory,一种两阶段选择性个性化框架:第一阶段查询请求与公共上下文缺失的用户特定信息,并检索覆盖全面的候选池;第二阶段根据响应级增量实用性,选择性引入可溯源的记忆单元或不引入。通过结构化SFT初始化、降维分阶段GRPO预热及嵌套多样本联合GRPO,逐步训练查询生成与证据采纳策略。在Goodreads、Amazon Reviews和Reddit的4,500个受控与自然任务中,该方法持续优于随机和词法记忆使用,改进于语义检索,且在本地生成能力增强时仍保持与前沿大模型记忆管道竞争力。其证据采纳条件基于公共上下文充分性,实现有选择而非默认的个性化。
原文摘要 · Abstract (English)
Personalized generation systems retrieve user history by request--memory relevance and inject it into the model context. Yet relevant history may concern the wrong preference aspect, duplicate public information, or provide insufficient support. We argue that personal memory should be used only when it adds utility beyond a public-only response. We propose TRACE-Memory, a two-stage framework for selective personalization. Stage 1 queries for user-specific information missing from the request and public context, then retrieves a coverage-oriented candidate pool. Stage 2 admits a compact subset of source-traceable evidence units, or the empty set, according to response-level incremental utility. We progressively train the query-generation and evidence-admission policies through structured SFT initialization, reduced-space stage-wise GRPO warm-up, and nested multi-sample Joint GRPO. Across 4,500 Controlled and Natural tasks from Goodreads, Amazon Reviews, and Reddit, TRACE-Memory consistently outperforms random and lexical memory use, improves over semantic retrieval, remains competitive with frontier-LLM memory pipelines as local generator capacity increases, and conditions evidence admission on public-context sufficiency, supporting selective rather than default personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。