用智能代理提升时间序列数据质量,显著降低预测误差。
Empowering Time Series Forecasting with LLM-Agents
- 基于时间序列元信息构建数据清洁代理,自动优化数据质量。
- 在交通流量数据上平均降低6%预测误差,跨模型和时长均有效。
- 适合关注数据治理的工业界研究者,尤其重视数据而非模型结构者。
大型语言模型驱动的智能体已成为自动化机器学习(AutoML)系统中有效的规划工具。尽管现有AutoML方法多聚焦于特征工程与模型架构搜索,但近期研究表明,轻量级模型在时间序列预测中常能达到顶尖性能。这一观察促使我们探索以提升数据质量而非优化模型架构作为AutoML在时间序列上的潜在突破口。本文提出DCATS:面向时间序列的数据中心型智能体,利用时间序列附带的元信息进行数据清洗,同时优化预测性能。我们在大规模交通流量预测数据集上,使用四种时间序列预测模型评估DCATS。结果表明,DCATS在所有测试模型与预测时长下,平均实现6%的误差降低,凸显了数据驱动方法在时间序列AutoML中的巨大潜力。
原文摘要 · Abstract (English)
Large Language Model (LLM) powered agents have emerged as effective planners for Automated Machine Learning (AutoML) systems. While most existing AutoML approaches focus on automating feature engineering and model architecture search, recent studies in time series forecasting suggest that lightweight models can often achieve state-of-the-art performance. This observation led us to explore improving data quality, rather than model architecture, as a potentially fruitful direction for AutoML on time series data. We propose DCATS, a Data-Centric Agent for Time Series. DCATS leverages metadata accompanying time series to clean data while optimizing forecasting performance. We evaluated DCATS using four time series forecasting models on a large-scale traffic volume forecasting dataset. Results demonstrate that DCATS achieves an average 6% error reduction across all tested models and time horizons, highlighting the potential of data-centric approaches in AutoML for time series forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。