发现GPT-4o性能随时间和周期波动,影响研究可靠性
Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
- 用固定提示长期测试GPT-4o,每3小时运行10次任务
- 性能波动占总方差20%,呈现日周期与周周期叠加特征
- 提醒研究人员警惕模型输出的时变性,尤其在重复实验中
大型语言模型(LLMs)在研究中既作为工具也作为研究对象被广泛使用。当前多数工作假设,在固定条件(相同模型快照、超参数和提示)下,LLM的性能是时间不变的,即平均输出质量稳定;否则将影响结果的可靠性和可复现性。为检验该假设,我们对GPT-4o在固定条件下进行了纵向研究:每隔三小时对同一物理任务进行十次查询,持续约三个月。对所得时间序列进行谱分析(傅里叶分析),发现显著的周期性波动,占总方差约20%。观察到的周期模式与日周期和周周期的相互作用一致。这些发现挑战了时间不变性的假设,并对涉及LLM的研究具有重要意义。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in research as both tools and objects of study. Much of this work assumes that LLM performance under fixed conditions (identical model snapshot, hyperparameters, and prompt) is time-invariant, meaning that average output quality remains stable over time; otherwise, reliability and reproducibility would be compromised. To test the assumption of time invariance, we conducted a longitudinal study of GPT-4o's average performance under fixed conditions. The LLM was queried to solve the same physics task ten times every three hours over approximately three months. Spectral (Fourier) analysis of the resulting time series revealed substantial periodic variability, accounting for about 20% of total variance. The observed periodic patterns are consistent with interacting daily and weekly rhythms. These findings challenge the assumption of time invariance and carry important implications for research involving LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。