arXiv:2607.23647cs.LGcs.AI2026-07

分离用户长期偏好、短期意图和曝光行为,提升长时推荐准确性

CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation

论文配图:CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation
图 1 · 摘自论文原文
  • 用冻结语言模型提取物品与反馈的语义原子,分设长短记忆
  • 在电商、新闻、短视频场景中,长期价值提升6.1%~7.6%
  • 通过反事实删除验证解释可信性,适合需要可解释推荐的场景

大型语言模型(LLMs)能以自然语言总结异构用户证据,但现有方法常将持久偏好、临时意图和曝光诱导行为混为一谈,导致推荐受反馈循环影响:重复曝光被误认为偏好,即时点击主导延迟满足,流畅解释未必反映排序依据。我们提出一种模型无关的长时推荐框架,利用冻结多模态语言模型将物品内容与反馈转化为基于证据的语义原子,并分别维护短期、长期和曝光记忆。通过倾向性加权更新减少策略引起的曝光偏差,保守离线评判器在行为支持约束下重排候选项以促进延迟满足。解释仅依赖关键证据原子,并通过反事实删除检验。我们提供识别结果,并在类电商、类新闻、类短视频环境中评估。十次随机种子实验显示,相比最强基线,本方法在长期折扣价值上分别提升6.1%、7.6%、6.7%。二十次配对消融实验表明,移除倾向性校正(0.739±0.191)或保守支持正则化(0.523±0.234)后性能显著下降。冻结指令语言模型在保留的释义基准上,语义原子的NDCG超过TF-IDF两倍以上。

原文摘要 · Abstract (English)

Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse enduring preferences, transient intent, and exposure-induced behavior into one profile. This makes recommendation vulnerable to feedback loops: repeated exposure is mistaken for preference, immediate clicks dominate delayed satisfaction, and fluent explanations need not reflect the ranking decision. We propose our method, a model-agnostic framework for long-horizon recommendation. Our method uses a frozen multimodal language model to convert item content and feedback into evidence-grounded semantic atoms, then maintains separate short-term, long-term, and exposure memories. Propensity-weighted updates reduce policy-induced exposure bias, while a conservative offline critic reranks candidates for delayed satisfaction under a behavior-support constraint. Explanations use only influential evidence atoms and are checked by counterfactual deletion. We provide an identification result and evaluate the framework in e-commerce-like, news-like, and short-video-like environments. Across ten seeds, our method improves discounted long-term value over the strongest alternative by 6.1%, 7.6%, and 6.7%, respectively. Twenty-seed paired ablations show significant value drops after removing propensity correction (0.739 +/- 0.191) or conservative support regularization (0.523 +/- 0.234). A frozen instruction language model also more than doubles semantic-atom NDCG over TF-IDF on a held-out paraphrase benchmark.

长时推荐因果建模可解释推荐语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。