用大模型发现时间序列中的隐藏规律,挑战科学发现新范式。
Can Large Language Models Adequately Perform Symbolic Reasoning Over Time Series?
- 结合大模型与遗传编程,构建闭环符号推理系统。
- 在真实时间序列上验证,大模型需依赖领域知识提升表现。
- 适合对自动化科学发现感兴趣的科研人员。
从时间序列数据中挖掘隐藏的符号规律,这一追求可追溯至开普勒发现行星运动定律。尽管大语言模型在结构化推理任务中展现潜力,但其从时间序列中推断可解释、上下文对齐的符号结构的能力仍待深入探索。为此,我们提出SymbolBench,一个涵盖三大任务(多变量符号回归、布尔网络推断、因果发现)的综合性基准,评估真实世界时间序列上的符号推理能力。相比以往仅限于简单代数方程的研究,SymbolBench覆盖多种复杂度的符号形式。我们进一步提出统一框架,将大模型与遗传编程结合,形成闭环符号推理系统,其中大模型同时承担预测与评估角色。实验结果揭示当前模型的关键优劣,强调融合领域知识、上下文对齐与推理结构对提升大模型在自动科学发现中表现的重要性。
原文摘要 · Abstract (English)
Uncovering hidden symbolic laws from time series data, as an aspiration dating back to Kepler's discovery of planetary motion, remains a core challenge in scientific discovery and artificial intelligence. While Large Language Models show promise in structured reasoning tasks, their ability to infer interpretable, context-aligned symbolic structures from time series data is still underexplored. To systematically evaluate this capability, we introduce SymbolBench, a comprehensive benchmark designed to assess symbolic reasoning over real-world time series across three tasks: multivariate symbolic regression, Boolean network inference, and causal discovery. Unlike prior efforts limited to simple algebraic equations, SymbolBench spans a diverse set of symbolic forms with varying complexity. We further propose a unified framework that integrates LLMs with genetic programming to form a closed-loop symbolic reasoning system, where LLMs act both as predictors and evaluators. Our empirical results reveal key strengths and limitations of current models, highlighting the importance of combining domain knowledge, context alignment, and reasoning structure to improve LLMs in automated scientific discovery. https://github.com/nuuuh/SymbolBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。