arXiv:2604.14972cs.IR2026-04被引 3

让推荐智能体学会自我改进决策逻辑,个性化推理而非仅记忆偏好。

SAGER: Self-Evolving User Policy Skills for Recommendation Agent

论文配图:SAGER: Self-Evolving User Policy Skills for Recommendation Agent
图 1 · 摘自论文原文
  • 为每位用户生成可演化的行为策略文档,实现个性化决策逻辑
  • 在4个公开数据集上超越现有方法,且性能提升与记忆积累无关
  • 适合研究个性化推荐系统、AI代理自我优化的学者和开发者

基于大语言模型的推荐智能体通过演化的用户语义记忆实现个性化,但其推理逻辑仍使用全用户通用的静态系统提示。这种不对称性导致推荐失败时,智能体只更新偏好记忆,却从不反思错误决策的逻辑,使推理结构始终不变。为此,我们提出SAGER(自演化推荐智能体),首个为每位用户配备专属策略技能的推荐框架。该技能是以自然语言结构化记录的个性化决策原则,随交互持续演化。SAGER采用双表示技能架构,将丰富演化能力与轻量推理注入分离;设计增量对比思维链引擎,通过对比被接受与未选物品诊断推理缺陷,同时保留已有先验;引入技能增强的列表级推理,建立细粒度决策边界,使演化技能产生实质性判别价值。在四个公开基准上的实验表明,SAGER达到当前最优性能,且增益与记忆积累正交,证实个性化推理过程是推荐效果提升的质变来源。

原文摘要 · Abstract (English)

Large language model (LLM) based recommendation agents personalize what they know through evolving per-user semantic memory, yet how they reason remains a universal, static system prompt shared identically across all users. This asymmetry is a fundamental bottleneck: when a recommendation fails, the agent updates its memory of user preferences but never interrogates the decision logic that produced the failure, leaving its reasoning process structurally unchanged regardless of how many mistakes it accumulates. To address this bottleneck, we propose SAGER (Self-Evolving Agent for Personalized Recommendation), the first recommendation agent framework in which each user is equipped with a dedicated policy skill, a structured natural-language document encoding personalized decision principles that evolves continuously through interaction. SAGER introduces a two-representation skill architecture that decouples a rich evolution substrate from a minimal inference-time injection, an incremental contrastive chain-of-thought engine that diagnoses reasoning flaws by contrasting accepted against unchosen items while preserving accumulated priors, and skill-augmented listwise reasoning that creates fine-grained decision boundaries where the evolved skill provides genuine discriminative value. Experiments on four public benchmarks demonstrate that SAGER achieves state-of-the-art performance, with gains orthogonal to memory accumulation, confirming that personalizing the reasoning process itself is a qualitatively distinct source of recommendation improvement.

推荐系统智能体自演化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。