提出因果基准评分法,评估时序模型在噪声下的鲁棒性。
Creating a Causally Grounded Rating Method for Assessing the Robustness of AI Models for Time-Series Forecasting
- 基于因果分析构建评分框架,检测输入扰动下的统计与混淆偏差。
- 多模态和时序专用模型在12种数据分布下表现更稳健且准确。
- 无需模型权重即可帮助用户直观比较不同模型的鲁棒性。
AI模型(包括时序专用和通用基础模型)在金融等领域展现出强大的时序预测能力,但对输入扰动高度敏感,易导致预测错误并削弱投资者与分析师的信任。为此,本文提出一种因果基准评分框架,通过分析多种噪声和错误输入场景下的统计偏差与混淆偏差,系统评估模型鲁棒性。实验涵盖多个行业股票数据,测试了6类输入扰动和12种数据分布,覆盖单模态与多模态模型(如基于视觉变换器的模型和基础模型)。结果表明,多模态及时间序列专用基础模型在鲁棒性和准确性上优于通用模型。进一步用户研究验证了该评分框架的有效性:结合预测误差与评分,显著降低用户比较不同模型鲁棒性的难度。研究成果使利益相关方在无模型权重与训练数据的情况下,也能理解模型行为,辅助决策。
原文摘要 · Abstract (English)
AI models, including both time-series-specific and general-purpose Foundation Models (FMs), have demonstrated strong potential in time-series forecasting across sectors like finance. However, these models are highly sensitive to input perturbations, which can lead to prediction errors and undermine trust among stakeholders, including investors and analysts. To address this challenge, we propose a causally grounded rating framework to systematically evaluate model robustness by analyzing statistical and confounding biases under various noisy and erroneous input scenarios. Our framework is applied to a large-scale experimental setup involving stock price data from multiple industries and evaluates both uni-modal and multi-modal models, including Vision Transformer-based (ViT) models and FMs. We introduce six types of input perturbations and twelve data distributions to assess model performance. Results indicate that multi-modal and time-series-specific FMs demonstrate greater robustness and accuracy compared to general-purpose models. Further, to validate our framework's usability, we conduct a user study showcasing time-series models' prediction errors along with our computed ratings. The study confirms that our ratings reduce the difficulty for users in comparing the robustness of different models. Our findings can help stakeholders understand model behaviors in terms of robustness and accuracy for better decision-making even without access to the model weights and training data, i.e., black-box settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。