测试大模型预测用户重复行为时间间隔能力,发现过多上下文反而降低效果。
Is More Context Always Better? Examining LLM Reasoning Capability for Time Interval Prediction
- 用零样本测试大模型在重复购买场景中的时间预测能力。
- 大模型表现不如专用机器学习模型,且过度增加上下文会降低准确率。
- 适合关注时序推理局限性与混合模型设计的研究者阅读。
大型语言模型(LLMs)在多个领域展现出出色的推理与预测能力,但其从结构化行为数据中推断时间规律的能力仍不明确。本文系统研究了大模型预测用户重复行为时间间隔(如重复购买)的能力,以及不同层次上下文信息对其预测行为的影响。基于一个简单但具有代表性的重复购买场景,我们在零样本设置下对比了先进大模型与统计及机器学习模型的性能。关键发现包括:第一,尽管大模型优于轻量级统计基线,但始终逊于专用机器学习模型,表明其捕捉定量时间结构的能力有限;第二,适度上下文可提升大模型准确性,但增加更多用户级细节反而导致性能下降。这一结果挑战了‘更多上下文带来更好推理’的假设。研究揭示了当前大模型在结构化时序推理中的根本局限,并为未来融合统计精度与语言灵活性的上下文感知混合模型设计提供指导。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning and prediction across different domains. Yet, their ability to infer temporal regularities from structured behavioral data remains underexplored. This paper presents a systematic study investigating whether LLMs can predict time intervals between recurring user actions, such as repeated purchases, and how different levels of contextual information shape their predictive behavior. Using a simple but representative repurchase scenario, we benchmark state-of-the-art LLMs in zero-shot settings against both statistical and machine-learning models. Two key findings emerge. First, while LLMs surpass lightweight statistical baselines, they consistently underperform dedicated machine-learning models, showing their limited ability to capture quantitative temporal structure. Second, although moderate context can improve LLM accuracy, adding further user-level detail degrades performance. These results challenge the assumption that "more context leads to better reasoning". Our study highlights fundamental limitations of today's LLMs in structured temporal inference and offers guidance for designing future context-aware hybrid models that integrate statistical precision with linguistic flexibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。