用可模拟性评估时间序列预测的自然语言解释,发现数值推理比模型大小更重要。
XForecast: Evaluating Natural Language Explanations for Time Series Forecasting
- 提出基于可模拟性的新评估指标,衡量人类能否根据解释复现预测结果。
- 实验表明该指标能有效区分优质与劣质解释,且与人工判断一致。
- 发现大模型生成解释的质量关键在数值推理能力,而非模型规模。
时间序列预测对决策至关重要,但理解模型预测依赖可解释性。传统可解释AI方法(如特征重要性)需专业知识,而自然语言解释(NLEs)更易被非专业人士理解。然而,由于时间序列数据中复杂的因果关系,评估预测解释仍具挑战。为此,本文提出两个基于可模拟性的新评估指标:评估人类代理是否能仅凭解释准确预测模型输出。实验表明,这些指标能有效区分高质量与低质量解释,并与人工判断高度一致。利用这些指标,进一步评估了当前主流大语言模型(LLMs)生成时间序列解释的能力,发现数值推理能力是影响解释质量的核心因素,而非模型规模。
原文摘要 · Abstract (English)
Time series forecasting aids decision-making, especially for stakeholders who rely on accurate predictions, making it very important to understand and explain these models to ensure informed decisions. Traditional explainable AI (XAI) methods, which underline feature or temporal importance, often require expert knowledge. In contrast, natural language explanations (NLEs) are more accessible to laypeople. However, evaluating forecast NLEs is difficult due to the complex causal relationships in time series data. To address this, we introduce two new performance metrics based on simulatability, assessing how well a human surrogate can predict model forecasts using the explanations. Experiments show these metrics differentiate good from poor explanations and align with human judgments. Utilizing these metrics, we further evaluate the ability of state-of-the-art large language models (LLMs) to generate explanations for time series data, finding that numerical reasoning, rather than model size, is the main factor influencing explanation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。