修复大模型推荐中因上下文漂移导致的性能下降问题
A Reproducibility Analysis of PO4ISR: Diagnosing and Mitigating Semantic Drift in LLM-Based Session Recommendation

- 用动态反思提示增强原模型的跨域适应能力
- 在Games和Bundle数据集上性能提升最高达96%
- 适合关注大模型可复现性与推荐系统稳定性的研究者
基于推理的大语言模型(如PO4ISR)在会话推荐任务中取得了新基准表现,但其在不同语义领域中的可复现性尚未被探索。本文对PO4ISR进行严谨的可复现性分析,揭示其在长会话中因标准推理提示导致严重上下文漂移,进而造成在语义复杂数据集(如Games和Bundle)上性能下降。为量化并解决这一稳定性差距,我们提出PO4ISR++,通过引入反思式提示与一致排名检测机制,实现对跨域线索的动态适应。在ML-1M、Games和Bundle上的实验表明,尽管原始模型在新领域表现不佳,而我们的改进版本恢复了性能,在Games上最高提升54%,在Bundle上最高提升96%。我们开源了复现基线与增强框架,以支持未来大模型推荐的可靠研究。
原文摘要 · Abstract (English)
Reasoning-based Large Language Models (LLMs) like PO4ISR have set new benchmarks in session-based recommendation. However, the reproducibility of their reasoning capabilities across diverse semantic domains remains unexplored. In this work, we conduct a rigorous reproducibility study of PO4ISR to assess its generalization limits. Our analysis reveals a critical failure mode: standard reasoning prompts suffer from severe contextual drift in long sessions, leading to performance degradation on semantically complex datasets like Games and Bundle. To quantify and resolve this stability gap, we introduce PO4ISR++, a robustness-enhanced implementation that integrates reflexive prompting and consistent rank detection. Unlike the original static prompting strategy, our approach dynamically adapts to cross-domain cues. We benchmark both the original implementation and our robust variant on ML-1M, Games, and Bundle. Our results confirm that while the original model struggles in new domains, our reproducible extension restores performance, yielding a stabilized gain of up to 54% on Games and 96% on Bundle. We release open-source artifacts, including the reproduced baseline and our enhanced framework, to facilitate reliable future research in LLM-based recommendation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。